Staff ML Engineer — Production Inference & Serving

Orbifold AI, Inc.

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Orbifold AI is building the infrastructure layer for Physical AI in a scale-oriented environment. You will own the serving architecture for production models, designing and running the end-to-end inference stack on Ray Serve and PyTorch across a heterogeneous GPU fleet with variable workloads.

We expect 3+ years building production ML systems, deep Python and PyTorch expertise, and a strong grasp of GPU execution, batching, quantization, and kernel-level optimization.

Qualifications

  • 3+ years building production ML systems with Python and PyTorch.
  • Proven ownership of inference performance for real workloads with measurable metrics.
  • Mental model of GPU execution: memory bandwidth, occupancy and kernel launch costs.
  • Experience operating distributed serving or compute frameworks under real load.
  • Comfort being measured on throughput, latency and cost over time.
  • Judgment on when to optimize and when to leave things as is.

Responsibilities

  • Own model serving end to end; design and run the serving layer for production models on Ray Serve and PyTorch.
  • Optimize throughput and utilization: batching, scheduling, concurrency, queueing.
  • Make models fit: quantization, memory layout, activation and cache management with verification.
  • Delve into compiler paths, custom kernels where beneficial and know when not to.
  • Handle high-volume video and multimodal inference workloads.
  • Build benchmarking harnesses to ensure reproducible performance claims.
  • Operate in production with autoscaling, fault tolerance, observability and cost per workload.

Skills

Python
PyTorch
Inference performance
GPU execution model
Distributed serving (Ray/Kubernetes)
Throughput & latency focus
Profiling & optimization

Tools

Ray Serve
TensorRT
Triton
XLA
Torch.compile
CUDA
TensorRT-LLM
Open-source serving stacks

Job description

Orbifold AI is building the infrastructure layer for Physical AI in a scale-oriented environment. You will own the serving architecture for production models, designing and running the end-to-end inference stack on Ray Serve and PyTorch across a heterogeneous GPU fleet with variable workloads.

We expect 3+ years building production ML systems, deep Python and PyTorch expertise, and a strong grasp of GPU execution, batching, quantization, and kernel-level optimization.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Inference & Performance Engineer
ML Inference & Performance Engineer

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Member of Technical Staff, ML Engineer (Inference & Performance)
Member of Technical Staff, ML Engineer (Inference & Performance)

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 250,000
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Inference Engineer - High-Performance AI Systems
ML Inference Engineer - High-Performance AI Systems

Together Computer Inc • San Francisco (CA)

On-site
USD 200,000 - 300,000
Startup equity
Health insurance
Competitive compensation
Senior ML Inference Engineer — Scale Production APIs
Senior ML Inference Engineer — Scale Production APIs

AssemblyAI, Inc. • New York (NY)

On-site
USD 190,000 - 225,000
Senior ML Inference Engineer — Production Systems
Senior ML Inference Engineer — Production Systems

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Remote AI/ML Engineer — Production Inference & RL
Remote AI/ML Engineer — Production Inference & RL

Boundless • Northern (KY)

Hybrid
USD 175,000 - 250,000
Competitive salary + equity
Health/dental/vision
Flexible PTO
+2
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Archetype AI Inc. • San Mateo (CA)

On-site
USD 180,000 - 280,000
Member of Technical Staff, Inference & Serving
Member of Technical Staff, Inference & Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000