Machine Learning Engineer, Applied AI Infrastructure

Bonfirevc

Palo Alto (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Orbifold AI in Palo Alto, CA is hiring a Machine Learning Engineer to scale and optimize the ML infrastructure behind multimodal video, training, and RL pipelines. You will own the systems that turn raw multimodal data into the training, evaluation, and RL signals our partners depend on, building scalable, fault-tolerant pipelines with PyTorch and Ray.

You will collaborate with research, data, and product engineering teams to translate modeling constraints into scalable infrastructure solutions,

Qualifications

  • 3+ years of software engineering experience with backend, distributed systems, or ML infrastructure.
  • Strong proficiency in Python and production-grade code.
  • Deep practical knowledge of PyTorch — including model serving, data loading bottlenecks, and memory management.

Responsibilities

  • Architect, build, and optimize distributed ML pipelines on Ray and PyTorch for multimodal data at scale.
  • Profile and tune training jobs and inference deployments to maximize GPU/CPU utilization and reduce latency.
  • Build abstractions and tools to deploy PyTorch models onto Ray clusters seamlessly.
  • Design and maintain high-throughput video processing pipelines feeding training, evaluation, and curation workloads.
  • Ensure high availability, fault tolerance, and observability of distributed compute systems.

Skills

Python
PyTorch
Ray
Distributed systems
Backend engineering
Problem solving

Tools

Docker
Kubernetes

Job description

Machine Learning Engineer, Applied AI Infrastructure

Palo Alto, CA (On-site)

Scale our Ray + PyTorch infrastructure for the multimodal video, training, and RL pipelines that power frontier robotics and world model teams. 3+ yrs distributed systems / ML infra.

About Orbifold AI

Orbifold AI is building the foundational infrastructure that the next generation of physical AI runs on. We work directly with leading robotics and world model research teams. Our work spans evaluation, model training, reinforcement learning, and the multimodal data systems that fuel them — one integrated research loop.

The bottleneck for physical AI is no longer model scale or computation. It is whether evaluation, training, and data can close the loop tightly enough to drive real progress. That loop is itself the infrastructure the next generation of physical AI will stand on, and it is what we are building.

Role Overview

We are hiring a Machine Learning Engineer to scale and optimize the ML infrastructure behind our pipelines. We process massive volumes of multimodal data — video, image, sensor, action — for some of the most demanding physical AI and world model teams in the field. Our foundation is built on PyTorch and Ray.

You will own the systems that turn raw multimodal data into the training, evaluation, and RL signals our partners depend on. Your work is the bridge between our research and our distributed compute infrastructure: making the pipelines performant, fault-tolerant, and ready to scale to the next order of magnitude.

This is highly applied infrastructure work with direct impact on what our partner models can do in the real world.

What You Will Work On
  • Architect, build, and optimize distributed ML pipelines on Ray (Ray Core, Ray Train, Ray Serve) and PyTorch, designed for the demands of multimodal video, image, and sensor data at scale
  • Profile and tune distributed training jobs and inference deployments to maximize GPU/CPU utilization and reduce latency
  • Build robust abstractions and internal tools that let our researchers and product engineers deploy PyTorch models onto our Ray clusters seamlessly
  • Design and maintain high-throughput video processing pipelines (e.g. FFmpeg, NVDEC/NVENC, frame-level indexing) that feed our curation, training, and evaluation workloads
  • Ensure the high availability, fault tolerance, and observability of our distributed compute systems
  • Build the serving infrastructure for our evaluation harnesses, verification models, and RL environments
  • Collaborate with research, data, and product engineering teams to translate modeling constraints into scalable infrastructure solutions
What We Are Looking For
  • 3+ years of software engineering experience with a strong focus on backend, distributed systems, or ML infrastructure
  • Strong proficiency in Python and production-grade code
  • Deep practical knowledge of PyTorch — including model serving, data loading bottlenecks, and memory management
  • Hands-on experience with Ray for scaling Python and machine learning applications
  • Solid understanding of distributed systems concepts: networking, concurrency, fault tolerance, parallel processing
  • Comfortable owning systems end to end in fast-paced applied research or startup environments
Nice to Have
  • Experience with large-scale video or multimodal data pipelines (e.g. FFmpeg, NVDEC/NVENC, 3D / point cloud handling)
  • Cloud-native infrastructure experience (Kubernetes, Docker) and major cloud providers (AWS, GCP, Azure)
  • Hardware accelerator experience (GPUs, TPUs) and low-level optimization (CUDA, C++)
  • Background in MLOps and automated CI/CD pipelines for machine learning
  • Familiarity with VLA models, world models, or robotics middleware (e.g. ROS/ROS2)
  • Experience with reinforcement learning environments or simulation infrastructure
Why This Role
  • Build the infrastructure that the next generation of physical AI will stand on
  • Work directly with the labs and companies shipping frontier robotics and world model systems
  • Own a critical layer of the stack end to end — from raw video and sensor ingest to distributed training and real-time evaluation serving
  • High ownership, fast iteration, and direct impact on deployed systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, Applied AI Infrastructure
Machine Learning Engineer, Applied AI Infrastructure

Orbifold AI Inc. • Palo Alto (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer, Physical AI
ML Systems Engineer, Physical AI

Orbifold AI • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer, Applied AI Infrastructure
Machine Learning Engineer, Applied AI Infrastructure

Orbifold AI • Palo Alto (CA)

On-site
USD 170,000 - 230,000
ML Infra Engineer — Ray + PyTorch Pipelines
ML Infra Engineer — Ray + PyTorch Pipelines

Orbifold AI • Palo Alto (CA)

On-site
USD 170,000 - 230,000
ML Infrastructure Engineer: Ray + PyTorch for Multimodal AI
ML Infrastructure Engineer: Ray + PyTorch for Multimodal AI

Orbifold AI Inc. • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Members of Technical Staff, Physical AI (Robotics / World Models)
Members of Technical Staff, Physical AI (Robotics / World Models)

Bonfirevc • Palo Alto (CA)

On-site
USD 100,000 - 150,000
ML Infra Engineer - Ray + PyTorch for Multimodal AI
ML Infra Engineer - Ray + PyTorch for Multimodal AI

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Members of Technical Staff, Physical AI (Robotics / World Models)
Members of Technical Staff, Physical AI (Robotics / World Models)

Orbifold AI • Palo Alto (CA)

On-site
USD 180,000 - 260,000
ML Infra Engineer: Scale Ray + PyTorch for Multimodal AI
ML Infra Engineer: Scale Ray + PyTorch for Multimodal AI

Orbifold AI • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000