ML Infrastructure Engineer

Maven Robotics, Inc.

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Maven Robotics, Inc. is seeking an exceptional Infrastructure Engineer to design and scale the backend systems that power machine learning. You will manage the core infrastructure used by AI and robotics teams, ensuring reliability in data management, compute workloads, and engineering workflows.

The ideal candidate will have experience in distributed systems, backend services, and GPU compute infrastructure. Strong programming skills in Python, Go, Rust, or C++ are essential, along with a proactive attitude in driving complex infrastructure projects.

Qualifications

  • Significant experience designing, building, and operating production backend infrastructure.
  • Strong programming ability in Python, Go, Rust, C++, or similar languages.
  • Experience with Kubernetes and distributed workload orchestration.

Responsibilities

  • Own the architecture and implementation of the machine learning infrastructure.
  • Build backend platforms for managing data and compute resources.
  • Design scalable systems for workload orchestration and storage.

Skills

Backend infrastructure design
Distributed systems
Data infrastructure management
GPUs and compute workload orchestration
Strong programming in Python, Go, Rust, or C++
Kubernetes
Infrastructure automation
Internal developer platforms

Job description

We are looking to recruit an exceptional Infrastructure Engineer to own and build the backend systems that power machine learning at Maven Robotics. In this role, you will design and scale the core infrastructure used by our AI and robotics teams to manage data, run compute workloads, store artifacts, monitor systems, and support rapidly growing engineering workflows.

You should be excited about distributed systems, backend services, data infrastructure, GPU compute, and high-reliability internal platforms. The ideal candidate has successfully built and operated similar systems before and can independently drive complex infrastructure projects from architecture through production operation. The underlying systems may be sophisticated, but the interfaces and workflows they expose should be reliable, intuitive, and easy for engineers to use.

Responsibilities
  • Own the architecture, implementation, reliability, and evolution of Maven’s machine learning infrastructure.
  • Build backend services and platforms for managing data, artifacts, jobs, logs, metadata, and compute resources across cloud and on-premise environments.
  • Design scalable systems for workload orchestration, storage, observability, security, and infrastructure automation.
  • Build intuitive internal tools and abstractions that make complex infrastructure easy for engineers to use.
  • Lead technical and commercial discussions with cloud and ML compute providers, including capacity planning, performance, reliability, and cost.
Qualifications
Must-have
  • Significant experience designing, building, and operating production backend, distributed, or compute infrastructure.
  • A track record of independently owning complex infrastructure projects from architecture through deployment and ongoing operation.
  • Strong programming ability in Python, Go, Rust, C++, or a similar backend or systems language.
  • Experience operating GPU compute infrastructure and orchestrating distributed workloads using Kubernetes, Ray, ZenML, or similar systems.
  • Experience designing and operating storage systems, observability platforms, infrastructure-as-code, and secure access controls.
  • Experience managing large-scale GPU fleets or hybrid cloud and on-premise compute environments.
  • Experience building internal developer platforms, CLIs, SDKs, or other self-service infrastructure tools.
  • Strong technical judgment, leadership, and communication skills, with the ability to drive decisions across teams and external partners.
  • Self-starter attitude with the ability to identify priorities and deliver durable solutions in a fast-paced startup environment.
Nice-to-have
  • Familiarity with GPU architecture, accelerator-aware software design, and profiling compute-intensive workloads.
  • Exposure to infrastructure supporting large-scale robot learning workloads, including policy training, simulation, and multimodal data pipelines.
  • Familiarity with SOC 2 controls, security practices, and audit readiness.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infra & Platform Engineer
Senior ML Infra & Platform Engineer

Maven Robotics, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior DevOps Engineer
Senior DevOps Engineer

Maven AGI • Boston (MA)

On-site
USD 140,000 - 210,000
Competitive salary
Comprehensive benefits
Meaningful equity stakes
+1
Senior DevOps Engineer
Senior DevOps Engineer

Maven AGI, Inc. • Boston (MA)

On-site
USD 150,000 - 230,000
Competitive salary
Comprehensive benefits
Equity stake
+1
ML Infrastructure Engineer, Training Redwood City, CA Fulltime
ML Infrastructure Engineer, Training Redwood City, CA Fulltime

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Machine Learning Ops & Infrastructure Engineer
Machine Learning Ops & Infrastructure Engineer

Noble Machines • Sunnyvale (CA)

On-site
USD 160,000 - 300,000
Backend Software Engineer (ML Infra)
Backend Software Engineer (ML Infra)

Rockstar • San Francisco (CA)

On-site
USD 100,000 - 130,000
Software Engineer, Machine Learning Infrastructure
Software Engineer, Machine Learning Infrastructure

Botauto • Houston (TX)

On-site
USD 90,000 - 120,000
Applied Machine Learning Platform Engineer
Applied Machine Learning Platform Engineer

Buzz Solutions • United States

On-site
USD 80,000 - 120,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Mach9 • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive salary
Health insurance
Flexible hours
+1
Sr. Platform Engineer, ML Infrastructure
Sr. Platform Engineer, ML Infrastructure

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000