Member of Technical Staff, ML Engineer (Inference & Performance)

Bonfirevc

Palo Alto (CA)

On-site

USD 180,000 - 250,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Orbifold AI in Palo Alto is seeking a Member of Technical Staff, ML Engineer (Inference & Performance) to own the end-to-end serving layer for production models on Ray Serve and PyTorch, across a heterogeneous GPU fleet. This role focuses on throughput, latency, memory fit, and cost efficiency.

You will push quantization, memory layout, and kernel-level optimizations, with opportunities to develop custom CUDA/Triton code, benchmarks, and production-grade observability for scalable, reliable

Qualifications

  • 3+ years building production ML systems with Python and PyTorch.
  • Experience owning inference performance for a real workload with measurable improvements.
  • Understanding of GPU execution: memory hierarchy, bandwidth, and kernel launch overhead.
  • Experience operating a distributed serving framework such as Ray or Kubernetes in production.
  • Comfortable being measured on throughput, latency, and cost.
  • Judgment about when to optimize and when to leave something alone.

Responsibilities

  • Own the serving end-to-end for production models on Ray Serve and PyTorch across a heterogeneous GPU fleet.
  • Design batching strategies, scheduling, concurrency, queueing, and related performance optimizations.
  • Implement quantization, memory layout, and cache management while validating output quality.
  • Develop compiler paths or custom CUDA/Triton kernels where beneficial.
  • Serve high-volume video and multimodal inference workloads with fast rollout cycles.
  • Build a benchmarking harness for reproducible performance claims and to catch regressions.
  • Operate in production with autoscaling, fault tolerance, and observability.
  • Onboard partner models onto our infrastructure and ensure parity with local results.

Skills

Python
PyTorch
Production ML systems
Ray
Kubernetes
GPU execution
Throughput optimization

Tools

Ray Serve
TensorRT-LLM
Torch.compile

Job description

Member of Technical Staff, ML Engineer (Inference & Performance)

Palo Alto, CA (On-site)

Own how fast, how cheaply and how reliably models run on our infrastructure. Throughput, latency and cost per unit of work, from the serving layer down to the kernel.

About Orbifold AI

Orbifold AI is building the infrastructure layer for Physical AI. As intelligent systems move beyond language into the physical world, they require a fundamentally new understanding of physics, action, and interaction.

We partner with leading robotics and world model research teams to advance the foundations of embodied intelligence, enabling intelligent systems to perceive, understand, and operate in the real world.

The standards we set, and the infrastructure we build to scale them, will define the next frontier of robotics and Physical AI.

Role Overview

Everything we deliver is produced by a model running on our infrastructure. Perception models, evaluation models, verification models, and increasingly our partners’ own models. The inputs are video and sensor data rather than short text, the volume grows every month, and the economics of the entire platform run through how efficiently that work executes.

You own that execution layer. Serving architecture on Ray and PyTorch, batching and scheduling, memory fit, quantization, compiler and kernel level work where it pays, and the profiling discipline that tells you which of those is actually the bottleneck today rather than which one is the most interesting.

This is a performance role with a direct product consequence. Every improvement you land buys more thorough evaluation, more verification passes, and more partner workloads on the same hardware. It is also a founding seat: the inference function at Orbifold is what you decide it is.

What You Will Work On

  • Own model serving end to end. Design and run the serving layer for our production models on Ray Serve and PyTorch, across a heterogeneous GPU fleet with very uneven workload shapes.
  • Drive throughput and utilization. Batching strategies, scheduling, concurrency, queueing, and the unglamorous work of finding out why a GPU is sitting at forty percent.
  • Make models fit. Quantization, precision selection, memory layout, activation and cache management, together with the measurement that proves output quality did not move when you did.
  • Go down a level when it pays. Compiler paths, operator fusion, custom CUDA or Triton kernels where the off-the-shelf version is leaving real performance on the table, and the judgment to know when it is not.
  • Serve the awkward workloads. High-volume video and multimodal inference, evaluation harnesses, and reinforcement learning environments that need many fast rollouts rather than one large request.
  • Build the benchmarking harness that makes performance claims reproducible instead of anecdotal, and keeps regressions from shipping quietly.
  • Run it in production. Autoscaling, fault tolerance, graceful degradation, observability, and cost per workload that anyone in the company can look up.
  • Bring partner models onto our infrastructure: packaging, throughput and memory fit, validation that what we return matches what they get locally, and an honest account of where it does not.

What We Are Looking For

  • 3+ years building production machine learning systems, with deep Python and PyTorch.
  • You have owned inference performance for a real workload, and you can describe the profile before and after in numbers rather than adjectives.
  • A working mental model of GPU execution: memory hierarchy and bandwidth, kernel launch overhead, occupancy, and where the time actually goes.
  • Experience operating a distributed serving or compute framework such as Ray or Kubernetes under real production load.
  • Comfortable being measured on throughput, latency and cost, including the weeks when the number does not move.
  • Judgment about when to optimize and when to leave something alone. Most of the value is in choosing correctly.

Nice to Have

  • CUDA, Triton, or compiler level work (torch.compile, TensorRT, XLA, or similar).
  • Video or multimodal inference at scale: hardware decode, preprocessing, batching across variable-length inputs.
  • Quantization, distillation or speculative methods taken all the way into production.
  • Ray Serve, vLLM, TensorRT-LLM, or comparable serving stacks.
  • Serving diffusion, video generation, or vision language action models.
  • Open-source contributions to a serving or performance stack that other engineers rely on.

Why This Role

  • Performance is the business. Compute is the dominant cost of what we do, so the work you own is visible at the level of what the company can afford to attempt.
  • The playbook is not written. Serving video and embodied models behaves almost nothing like serving text, and very few people have solved it at scale yet.
  • Full vertical ownership, from the serving API down to the kernel, without an abstraction layer separating you from the hardware.
  • Frontier workloads and real hardware, at a size where a single good decision is measurable within a week.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Systems Engineer, Physical AI
ML Systems Engineer, Physical AI

Orbifold AI • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff, Platform Engineer (Console & SDK)
Member of Technical Staff, Platform Engineer (Console & SDK)

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Founding Inference Engineer
Founding Inference Engineer

General Compute • San Francisco (CA)

On-site
USD 180,000 - 320,000
Member of Technical Staff, Platform Engineer (Console & SDK)
Member of Technical Staff, Platform Engineer (Console & SDK)

Orbifold AI, Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Member of Technical Staff, AI Compute & Data Infrastructure
Member of Technical Staff, AI Compute & Data Infrastructure

Vinci • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
ML Inference & Performance Engineer
ML Inference & Performance Engineer

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda Innovation • Northern (KY)

On-site
USD 150,000 - 230,000
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1