ML Infrastructure Engineer

Zipline

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Zipline is shaping autonomous systems by building scalable data and ML infrastructure. You will own end-to-end data pipelines, datasets, and model deployment to accelerate learning and delivery in real-world robotics environments.

The role emphasizes reproducible ML workflows, production-grade Python, and collaboration across cloud and on‑prem systems to support mission-critical operations.

Qualifications

  • Experience building reproducible data and ML pipelines.
  • Knowledge of datasets, training, evaluation, optimization and modern DL workflows.
  • 3+ years software engineering, including ML infrastructure or safety-critical environments.
  • Strong Python production practices; design of APIs, services, and workflows.
  • Experience with Kubernetes for production workloads.
  • Cloud/on‑prem infrastructure experience (AWS preferred) and IaC tools.
  • Proficiency with PyTorch or similar ML frameworks.
  • Ownership, communication, and secure systems mindset.
  • Willingness to work across cloud, data platforms, tooling, and robotics constraints.
  • Experience monitoring data stats, performance, pipeline failures and signals.
  • Experience with annotation systems or active-learning tools.
  • Experience with large-scale training systems, feature stores, model registries, or experiment tracking.
  • Experience deploying ML systems on real robots or autonomous hardware.

Responsibilities

  • Build and operate software infrastructure enabling learning over large-scale data.
  • Design scalable data and ML infra for autonomy teams including dataset handling and deployment.
  • Own and improve data pipelines feeding the ML development loop.
  • Identify bottlenecks in ML cycles, focusing on orchestration, performance, and reproducibility.

Skills

Python production
Kubernetes
AWS Cloud
ML Pipelines
Datasets
Model Deployment
Telemetry Monitoring
Security & Reliability
Data Platforms
Experiment Tracking

Tools

Terraform
CloudFormation
PyTorch
Git
APIs

Job description

  • Build and operate software infrastructure that enables learning algorithms to leverage Zipline’s large-scale (quickly growing!) fleet data
  • Design scalable, maintainable data and ML infrastructure for autonomy teams, including dataset creation, validation, training, evaluation, and deployment
  • Own and improve data pipelines that feed into the ML development loop
  • Identify and mitigate bottlenecks in the ML development cycle, especially around orchestration, performance, and reproducibility to increase the rate at which we can improve and scale the delivery experience

Experience building reproducible data pipelines and machine-learning pipelinesWorking knowledge of ML concepts such as datasets, training, evaluation, optimization, statistics, and modern deep learning workflows3+ years of professional software engineering experience, ideally including ML infrastructure, data infrastructure, robotics, autonomy, aerospace, medical devices, or another safety-critical hardware/product environmentStrong software engineering practices in Python in a production setting; comfort designing APIs, services, schemas, jobs, and operational workflowsExperience with Kubernetes or other container orchestration systems for production workloadsExperience with cloud and on-premise production infrastructure, preferably AWS, and infrastructure-as-code tools such as Terraform or CloudFormationExperience with PyTorch or similar ML frameworksStrong ownership, clear communication, and interest in building secure systems for mission-critical workflowsGeneralist mindset and willingness to work across cloud services, data platforms, developer tooling, and embedded/robotics-adjacent constraintsExperience monitoring data statistics, system performance metrics, pipeline failures, and model/evaluation signalsExperience with annotation systems, dataset inspection tooling, or active-learning workflowsExperience with large-scale training systems, feature stores, data/versioned artifact stores, model registries, or experiment trackingExperience deploying or evaluating ML systems on real robots, autonomous vehicles, drones, or other hardware products

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer for Scalable Autonomous Systems
ML Infrastructure Engineer for Scalable Autonomous Systems

Zipline • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Zipline • South San Francisco (CA)

On-site
USD 190,000 - 250,000
Equity compensation
Discretionary bonuses
Medical, dental, vision insurance
+1
ML Infrastructure Engineer — Scale ML for Drones (Equity)
ML Infrastructure Engineer — Scale ML for Drones (Equity)

Zipline • South San Francisco (CA)

On-site
USD 190,000 - 250,000
Equity compensation
Discretionary bonuses
Medical, dental, vision insurance
+1
Staff Machine Learning Systems & Reliability Engineer (Moveworks)
Staff Machine Learning Systems & Reliability Engineer (Moveworks)

ServiceNow • Mountain View (CA)

On-site
USD 250,000 - 320,000
Generous family leave
Annual learning stipend
Flexible PTO
+2
ML Infrastructure Engineer
ML Infrastructure Engineer

Mach9 Robotics Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

Alexander Chapman Ltd • New York (NY)

On-site
USD 120,000 - 170,000
Health insurance
Competitive equity
ML Infra Engineer (Data Systems)
ML Infra Engineer (Data Systems)

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

On-site
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2