ML Infrastructure Engineer: Scale Ray + PyTorch Pipelines

Bonfirevc

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Orbifold AI in Palo Alto builds the foundational infrastructure that the next generation of physical AI runs on. You will scale and optimize ML infrastructure behind pipelines that process multimodal data—video, images, sensors—for world-model and robotics teams, using Ray and PyTorch.

You will own end-to-end systems, focus on performance, fault tolerance, and observability, and collaborate with research, data, and product engineers to translate modeling constraints into scalable, reliable

Qualifications

  • 3+ years of software engineering experience with a focus on backend, distributed systems, or ML infrastructure.
  • Strong proficiency in Python and production-grade code.
  • Deep practical knowledge of PyTorch - including model serving, data loading bottlenecks, and memory management.
  • Hands-on experience with Ray for scaling Python and ML applications.
  • Solid understanding of distributed systems concepts: networking, concurrency, fault tolerance, parallel processing.
  • Comfortable owning systems end to end in fast-paced applied research or startup environments.

Responsibilities

  • Architect, build, and optimize distributed ML pipelines on Ray (Ray Core, Ray Train, Ray Serve) and PyTorch, designed for multimodal data at scale.
  • Profile and tune distributed training jobs and inference deployments to maximize GPU/CPU utilization and reduce latency.
  • Build robust abstractions and internal tools to deploy PyTorch models onto Ray clusters seamlessly.
  • Design and maintain high-throughput video processing pipelines feeding training and evaluation workloads.
  • Ensure high availability, fault tolerance, and observability of distributed compute systems.
  • Collaborate with research, data, and product engineering teams to translate modeling constraints into scalable infrastructure solutions.

Skills

Python
PyTorch
Distributed systems
ML infra
Backend engineering
Research collaboration

Tools

Ray
PyTorch

Job description

Orbifold AI in Palo Alto builds the foundational infrastructure that the next generation of physical AI runs on. You will scale and optimize ML infrastructure behind pipelines that process multimodal data—video, images, sensors—for world-model and robotics teams, using Ray and PyTorch.

You will own end-to-end systems, focus on performance, fault tolerance, and observability, and collaborate with research, data, and product engineers to translate modeling constraints into scalable, reliable

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer — Ray + PyTorch Pipelines
ML Infra Engineer — Ray + PyTorch Pipelines

Orbifold AI • Palo Alto (CA)

On-site
USD 170,000 - 230,000
ML Infra Engineer: Scale Ray + PyTorch for Multimodal AI
ML Infra Engineer: Scale Ray + PyTorch for Multimodal AI

Orbifold AI • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)

Orbifold AI • Palo Alto (CA)

On-site
USD 170,000 - 230,000
ML Systems Engineer, Physical AI
ML Systems Engineer, Physical AI

Orbifold AI • Palo Alto (CA)

On-site
USD 120,000 - 160,000
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infra Engineer, Modeling
ML Infra Engineer, Modeling

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Architect — Real-Time AI Pipelines & Scale
ML Infra Architect — Real-Time AI Pipelines & Scale

Mach9 • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive salary
Health insurance
Flexible hours
+1