Software Engineer/Senior Software Engineer (Data & Machine Learning Platform)

PlusAI

Santa Clara (CA)

Hybrid

USD 150,000 - 190,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Unlimited PTO
Flexible working
Health benefits
Equity, 401k and FSA
Professional development
Catered lunches

Job summary

PlusAI is seeking a Software Engineer to strengthen our Kubernetes platform, focusing on reliability, onboarding, and scalable batch/workflow capabilities for petabyte-scale processing.

You will operate production Kubernetes clusters end to end, implement GitOps with Argo CD, Helm, and Kustomize, and collaborate on multi-tenant compute, GPU scheduling, and data-processing pipelines with Spark.

Qualifications

  • BS, MS, or PhD in Computer Science or related field, or equivalent practical experience.
  • Hands-on experience operating production Kubernetes clusters with node lifecycle, upgrades, troubleshooting, and GitOps and infrastructure-as-code.
  • Self-driven with ownership and ability to drive projects end to end.
  • Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management.

Responsibilities

  • Improve reliability and efficiency of Kubernetes infrastructure.
  • Make workload onboarding simpler and more self-service.
  • Build reusable batch and workflow capabilities for petabyte-scale processing.
  • Operate production Kubernetes clusters end to end including provisioning, control planes, GPUs, networking, and storage.
  • Develop GitOps-based delivery using Argo CD, Helm, and Kustomize.
  • Create multi-tenant platform capabilities for scheduling, isolation, storage, networking, and observability.
  • Build and enhance distributed batch/workflow platforms for Spark processing and GPU replay/simulation.

Skills

Kubernetes
Platform engineering
Distributed data processing
ML/GPU infrastructure
Multi-tenant compute systems
GitOps
Argo CD
Helm
Kustomize
Spark
Ray
Kubeflow
Argo Workflows
Delta Lake
Apache Iceberg
GPU workload scheduling
Autoscaling

Education

BS, MS, or PhD in Computer Science or related field

Tools

Argo CD
Helm
Kustomize
Kubeflow
Argo Workflows
Spark
Delta Lake
Apache Iceberg

Job description

  • All key offline workloads — large-scale data processing, simulation, auto-labeling, scenario mining, and model training — run on the compute platform this role owns
  • In this role, you will improve the reliability and efficiency of our Kubernetes infrastructure, make workload onboarding simpler and more self-service, and build reusable batch and workflow capabilities for petabyte-scale processing
  • Operate and evolve our production Kubernetes clusters end to end: bare-metal provisioning automation, highly available control planes, node lifecycle, GPU container runtime, networking, and storage
  • Build safe, repeatable GitOps-based delivery for platform services and user applications using tools such as Argo CD, Helm, and Kustomize
  • Develop shared multi-tenant platform capabilities for scheduling, resource isolation, storage, networking, access control, secrets, and observability while improving CPU/GPU utilization and cost efficiency
  • Build and improve reusable distributed batch and workflow platforms for Spark data processing and GPU-based replay and simulation
  • Ensure that your work is performed in accordance with the company’s Quality Management System (QMS) requirements and contribute to continuous improvement efforts
Benefits
  • Unlimited PTO: Unlimited paid time off on top of company holidays to recharge and enjoy as needed
  • Flexible working: Flexibile work arrangements and technology to help manage life’s demands at the office and remote
  • Health benefits: Tiered options in medical, dental, and vision insurance to fit your unique needs
  • Future building: Equity, 401k and FSA options for all eligible employees to help build for your future
  • Professional development: Company-sponsored professional development opportunities such as conferences, training, and workshops
  • Catered lunches: Daily catered lunches at both of our Santa Clara and Fremont offices

We are looking for strong Kubernetes and platform-engineering fundamentals, depth in at least one adjacent area—distributed data processing, ML/GPU infrastructure, or multi-tenant compute systems—and the curiosity and ownership to grow across the others

We are open to candidates at either the Software Engineer or Senior Software Engineer level. Level will be determined by experience, technical depth, scope of ownership, and demonstrated impact.

You do not need experience with every technology in our stack; we value strong fundamentals, ownership, and the ability to learn

BS, MS, or PhD in Computer Science or a related technical field, or equivalent practical experience

Hands-on experience operating production Kubernetes clusters — node lifecycle, upgrades, troubleshooting — plus GitOps and infrastructure-as-code experience

Self-driven with a strong sense of ownership: a quick learner who is eager to take responsibility and drive projects forward end to end

Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management

Experience with Ray or Kubeflow

Experience operating large-scale distributed data-processing and workflow systems, with hands-on depth in a system such as Apache Spark and working knowledge of Argo Workflows or an equivalent orchestrator

Experience with lakehouse technologies such as Delta Lake or Apache Iceberg

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer (SE / Sr SE), Data & ML Platform
Software Engineer (SE / Sr SE), Data & ML Platform

Plus 2 • Santa Clara (CA)

On-site
USD 150,000 - 200,000
Software Engineer/Senior Software Engineer, Data & ML Platform
Software Engineer/Senior Software Engineer, Data & ML Platform

PlusAI • Santa Clara (CA)

On-site
USD 135,000 - 200,000
Senior Software Engineer
Senior Software Engineer

CoreWeave • Livingston (NJ)

On-site
USD 140,000 - 200,000
Senior Software Engineer, Developer Platform
Senior Software Engineer, Developer Platform

Bot Auto • San Francisco (CA)

On-site
USD 180,000 - 250,000
Staff Software Engineer - Data Platform
Staff Software Engineer - Data Platform

Hive • San Francisco (CA)

On-site
Medical, dental and vision benefits plus FSA
Family leave
401(k) program
+3
Senior DevOps Engineer
Senior DevOps Engineer

Newton Research • Needham (MA)

On-site
USD 150,000 - 175,000
Equity
Autonomy over tech strategy
Competitive compensation
Senior DevOps Engineer
Senior DevOps Engineer

Newton Research • Boston (MA)

On-site
USD 150,000 - 175,000
Equity
Competitive salary
Benefits
Senior Platform Engineer, Data & ML Infrastructure
Senior Platform Engineer, Data & ML Infrastructure

PlusAI • Santa Clara (CA)

On-site
USD 135,000 - 200,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000
Principal Platform Engineer
Principal Platform Engineer

European Recruitment BV • United States

On-site
USD 150,000 - 200,000