Senior Inference Optimization ML Engineer

Rhoda AI

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Rhoda AI in Palo Alto is building the next generation of generalist intelligent robots. You will own end-to-end inference performance, diagnosing latency, throughput, and efficiency of large foundation models in production.

Collaborate with research engineers to translate model innovations into optimized, deployment-ready implementations using PyTorch, Triton, CUDA, and related tooling for multimodal models.

Qualifications

  • 3+ years of experience in inference optimization or ML systems.
  • Strong debugging and measurement skills to identify bottlenecks.
  • Experience with large-model deployments in production.

Responsibilities

  • Own inference performance end-to-end — diagnose and improve latency, throughput, and efficiency of large foundation models in production
  • Build systematic performance attribution: latency decomposition, bottleneck identification, and prioritization across model families
  • Apply and develop optimization techniques including quantization, pruning, distillation, operator fusion, and model compilation
  • Optimize attention mechanisms, KV caching, and memory layouts for large multimodal models
  • Work with kernel-level tooling to identify hotspots and implement or tune custom kernels where needed
  • Build benchmarking and regression detection infrastructure: latency baselines, throughput curves, and automated detection of performance regressions across model versions
  • Collaborate closely with research engineers to translate model innovations into optimized, deployment-ready implementations

Skills

Inference optimization
ML systems
Performance debugging
PyTorch
JAX

Tools

PyTorch
JAX
TensorRT
Triton
CUDA

Job description

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

What You\'ll Do
  • Own inference performance end-to-end — diagnose and improve latency, throughput, and efficiency of large foundation models in production

  • Build systematic performance attribution: latency decomposition (compute vs. memory bandwidth vs. I/O), bottleneck identification, and prioritization across model families

  • Apply and develop optimization techniques including quantization, pruning, distillation, operator fusion, and model compilation (e.g., TensorRT, torch.compile, XLA)

  • Optimize attention mechanisms, KV caching, and memory layouts for large multimodal models (vision, video, language, proprioception)

  • Work with kernel-level tooling (e.g., CUDA, Triton) to identify hotspots and implement or tune custom kernels where needed

  • Build benchmarking and regression detection infrastructure: latency baselines, throughput curves, and automated detection of performance regressions across model versions

  • Collaborate closely with research engineers to translate model innovations into optimized, deployment-ready implementations

What We\'re Looking For
  • 3+ years of experience in inference optimization, ML systems, or a closely related field

  • Deep hands-on experience with modern ML stacks (PyTorch required; JAX a plus)

  • Strong understanding of compute, memory bandwidth, and I/O bottlenecks in large model inference

  • Experience with model optimization techniques: quantization (INT8/FP8/AWQ), distillation, pruning, and compilation

  • Familiarity with inference serving frameworks (e.g., Triton, TensorRT, vLLM, TorchServe)

  • Exceptional debugging and measurement ability: turn "inference is slow" into clear bottlenecks, experiments, and validated improvements

  • High ownership mindset and comfort in a fast-moving environment

Nice to Have (But Not Required)
  • GPU kernel or compiler-level experience (CUDA, Triton, graph capture, operator fusion)

  • Experience with multimodal or video model inference (variable-length sequences, packing/bucketing)

  • Familiarity with edge/cloud hybrid deployment patterns and on-robot inference constraints

  • Experience with speculative decoding, continuous batching, or other LLM serving optimizations

  • Background in streaming or low-latency systems relevant to real-time robot control

Why This Role
  • Direct leverage on research velocity and real-world robot performance — every efficiency gain you make accelerates model iteration and tightens the loop between model and robot behavior

  • Own the optimization layer that determines how quickly and efficiently our foundation models run in the real world — high ownership, high impact, small elite team

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Optimization Engineer for Real-World Robotics
Inference Optimization Engineer for Real-World Robotics

Rhoda AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Member of Technical Staff- Efficient Modeling
Research Member of Technical Staff- Efficient Modeling

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000
Research Member of Technical Staff- Efficient Modeling
Research Member of Technical Staff- Efficient Modeling

Rhoda AI • Mountain View (WY)

On-site
USD 150,000 - 230,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 180,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 200,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Research Member of Technical Staff- Post-training & Robot Learning
Research Member of Technical Staff- Post-training & Robot Learning

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000