Senior Platform Engineer, Data & ML Infrastructure

PlusAI

Santa Clara (CA)

On-site

USD 135,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

PlusAI in Santa Clara, CA, is seeking a Software Engineer or Senior Software Engineer to strengthen our Kubernetes infrastructure, simplify workload onboarding, and create reusable batch and workflow capabilities for petabyte-scale processing.

You will work on production clusters, implement GitOps workflows with Argo CD, Helm, and Kustomize, and advance GPU/ML workload scheduling and multi-tenant resource management. A strong foundation in distributed systems is needed.

Qualifications

  • Hands-on experience operating production Kubernetes clusters and GitOps/infrastructure-as-code
  • Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling
  • Self-driven with ownership and strong learning ability

Responsibilities

  • Operate and evolve production Kubernetes clusters end to end (provisioning, control planes, node lifecycle, GPU container runtime)
  • Build safe, repeatable GitOps-based delivery for platform services and user apps using Argo CD, Helm, and Kustomize
  • Develop multi-tenant platform capabilities for scheduling, isolation, storage, networking, access control, secrets and observability
  • Build reusable distributed batch/workflow platforms for Spark processing and GPU-based replay/simulation
  • Adhere to QMS requirements and contribute to continuous improvement efforts

Skills

Kubernetes
GitOps
Ownership
GPU scheduling

Education

BS/MS/PhD in CS or related

Tools

Argo CD
Helm
Kustomize
Ray
Kubeflow
Spark
Argo Workflows
Delta Lake
Apache Iceberg

Job description

PlusAI in Santa Clara, CA, is seeking a Software Engineer or Senior Software Engineer to strengthen our Kubernetes infrastructure, simplify workload onboarding, and create reusable batch and workflow capabilities for petabyte-scale processing.

You will work on production clusters, implement GitOps workflows with Argo CD, Helm, and Kustomize, and advance GPU/ML workload scheduling and multi-tenant resource management. A strong foundation in distributed systems is needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kubernetes Platform Engineer, Data & ML
Kubernetes Platform Engineer, Data & ML

Plus 2 • Santa Clara (CA)

On-site
USD 150,000 - 200,000
Software Engineer/Senior Software Engineer, Data & ML Platform
Software Engineer/Senior Software Engineer, Data & ML Platform

PlusAI • Santa Clara (CA)

On-site
USD 135,000 - 200,000
Software Engineer (SE / Sr SE), Data & ML Platform
Software Engineer (SE / Sr SE), Data & ML Platform

Plus 2 • Santa Clara (CA)

On-site
USD 150,000 - 200,000
Senior Platform Engineer, Data & AI (Petabyte-Scale)
Senior Platform Engineer, Data & AI (Petabyte-Scale)

C3 AI • Redwood City (CA)

On-site
USD 145,000 - 187,000
Equity plan
Excellent benefits
Senior DevOps Platform Engineer - AI/ML at Scale (Equity)
Senior DevOps Platform Engineer - AI/ML at Scale (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior Platform Engineer - Cloud, AI & Distributed Systems
Senior Platform Engineer - Cloud, AI & Distributed Systems

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Platform & Infra Engineer (Kubernetes, GPUs)
Senior ML Platform & Infra Engineer (Kubernetes, GPUs)

IDR, Inc. • Los Angeles (CA)

On-site
USD 180,000 - 240,000
Senior Platform & Data Infrastructure Engineer
Senior Platform & Data Infrastructure Engineer

Thomas Talent Network • San Francisco (CA)

On-site
USD 250,000 - 350,000
Senior ML Platform Engineer - Cloud, Kubernetes & CI/CD
Senior ML Platform Engineer - Cloud, Kubernetes & CI/CD

Jobtailor • New York (NY)

On-site
USD 140,000 - 190,000
Principal Platform Engineer
Principal Platform Engineer

European Recruitment BV • United States

On-site
USD 150,000 - 200,000