Staff Software Engineer, HPC

Zoox

Foster City (CA)

On-site

USD 180,000 - 260,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zoox is seeking an experienced Staff Software Engineer to build and operate a high-performance computing (HPC) platform that scales with our autonomous vehicle development. You will modernize our HPC stack using Ray.io, SLURM, and Kubernetes, ensuring reliability and world-class developer velocity across teams from data engineering to AI model training.

You will define Zoox's HPC strategy, work with Autonomy and Software stakeholders, and lead cross-team initiatives to improve compute throughput

Qualifications

  • Experience designing and operating large-scale distributed systems in production.
  • Proficiency with Python and distributed compute workloads.
  • Experience with Ray.io and Kubernetes.
  • Strong scheduling/scaling knowledge and optimization mindset.
  • Cloud infrastructure experience (AWS or similar).
  • Ability to mentor and lead cross-functional teams.

Responsibilities

  • Design and implement core services for distributed compute infrastructure supporting hundreds of thousands of concurrent jobs.
  • Work with customer teams and other infrastructure groups to build a multi-year HPC roadmap.
  • Lead multi-quarter, cross-team initiatives driving org-wide improvements.
  • Create production-grade APIs, SDKs, and tools for Zoox engineers to run large-scale workloads.
  • Design and improve job scheduling algorithms and auto-scaling policies for reliability.
  • Develop multi-region orchestration strategies optimizing data locality and performance.
  • Identify and resolve systemic reliability and performance issues via profiling and collaboration.
  • Evaluate new technologies to enhance compute and storage capabilities.
  • Develop capacity planning tools and forecasting models for growing compute needs.
  • Mentor junior engineers and guide career development.

Skills

Distributed systems
Ray.io
Kubernetes
Python
AWS
Job scheduling
Capacity planning
Mentoring
Performance optimization
Orchestration

Tools

Ray Core
Ray Data
SLURM
Kubernetes

Job description

Zoox is looking for an experienced Staff Software Engineer to build, scale, and operate our custom High-Performance Computing infrastructure. As Zoox scales its autonomous vehicle development, our HPC platform must keep pace with rapidly growing compute, storage, and scheduling demands across the company. You will modernize our HPC platform—built on industry-leading technologies like Ray.io, SLURM, and Kubernetes—with a focus on reliability, scalability, and world-class developer velocity.

These HPC services form the backbone of development workflows across all Zoox software teams, from data engineering to training our AI models in Perception, Planner, Prediction, to Simulation, and more. You will have a direct impact on the productivity and effectiveness of every engineering team at Zoox.

The position comes with a high degree of independence and the opportunity to define Zoox's HPC platform strategy, both technically and organizationally. You will work closely with stakeholders in Autonomy and Software teams to understand their workload requirements and translate them into robust, scalable infrastructure.

  • Design and implement core services and abstractions for distributed compute infrastructure supporting hundreds of thousands of concurrent jobs
  • Work with customer teams and other infrastructure teams to build a multiyear software engineering roadmap for the HPC platform
  • Lead multi-quarter, cross team initiatives that drive org-wide improvements
  • Create production-grade APIs, SDKs, and tools that make it easy for engineers across Zoox to run large-scale distributed workloads
  • Design and improve job scheduling algorithms and auto-scaling policies to maximize reliability and resource availability
  • Design multi-region orchestration strategies that optimize for data locality, reliability, and performance
  • Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners across multiple teams
  • Evaluate new technologies and paradigms that improve Zoox's computational and storage capabilities
  • Develop capacity planning tools and forecasting models to support Zoox's growing compute needs
  • Mentor junior engineers, guiding them through their career development
  • Experience designing and operating large-scale distributed systems in production
  • Experience with Ray.io, particularly Ray Core and Ray Data (or equivalent technologies)
  • Experience with Kubernetes, particularly for heterogeneous workloads
  • Experience with cloud infrastructure on AWS or similar providers
  • Track record of shipping and operating reliable, highly availablescalable infrastructure
  • Demonstrated ability to prioritize development work and build cross-functional consensus around technical tradeoffs
  • Proficiency with Python
  • Exposure to machine learning workloads (training, inference, data generation)
  • Experience with Kubernetes or SLURM at scale (>10k+ nodes)
  • Experience with SLURM workload manager and advanced scheduling policies
  • Background in algorithmic optimization or operations research
  • Experience building developer tools and platforms used by large engineering organizations
About Zoox

Zoox is developing the first ground-up, fully autonomous vehicle fleet and the supporting ecosystem required to bring this technology to market. Sitting at the intersection of robotics, machine learning, and design, Zoox aims to provide the next generation of mobility-as-a-service in urban environments. We’re looking for top talent that shares our passion and wants to be part of a fast-moving and highly execution-oriented team.

Follow us on LinkedIn

A Final Note: You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, HPC
Staff Software Engineer, HPC

Zoox • Boston (MA)

On-site
USD 230,000 - 309,000
Staff Software Engineer, HPC Software
Staff Software Engineer, HPC Software

Front Door Defense • Foster City (CA), Northern (KY)

Hybrid
USD 230,000 - 295,000
Staff Software Engineer — Scalable HPC Platform Lead
Staff Software Engineer — Scalable HPC Platform Lead

Zoox • Foster City (CA)

On-site
USD 180,000 - 260,000
Staff HPC Platform Engineer for Distributed Compute
Staff HPC Platform Engineer for Distributed Compute

Zoox • Boston (MA)

On-site
USD 230,000 - 309,000
Staff HPC Engineer: Build Scalable Distributed Compute
Staff HPC Engineer: Build Scalable Distributed Compute

Front Door Defense • Foster City (CA), Northern (KY)

Hybrid
USD 230,000 - 295,000
Staff Software Engineer - Operational Tools
Staff Software Engineer - Operational Tools

Dormont Manufacturing Co • Foster City (CA)

On-site
USD 120,000 - 150,000
Senior Automation & Middleware Engineer
Senior Automation & Middleware Engineer

Zoox • Foster City (CA)

On-site
USD 195,000 - 240,000
Health insurance
Long-term care insurance
Long-term and short-term disability
+2
Staff Software Engineer, Operational Applications & Platforms
Staff Software Engineer, Operational Applications & Platforms

Front Door Defense • Foster City (CA)

On-site
USD 250,000 - 300,000
Senior/Staff Software Engineer - Embedded Linux Operating Systems
Senior/Staff Software Engineer - Embedded Linux Operating Systems

Zoox • Foster City (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - Full Stack Operational Tools
Senior Software Engineer - Full Stack Operational Tools

Zoox • Foster City (CA)

On-site
USD 210,000 - 305,000
Health insurance
Paid time off
Long-term care insurance
+1