Robotics AI Inference Optimization Engineer

Rhoda AI

Mountain View (WY)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Rhoda AI is hiring an Inference Optimization MLE to build and operate systems that make foundation models run fast in production. Youll squeeze performance across cloud and on-robot targets, collaborating with research to close the training-deployment gap.

Youll own end-to-end optimization, implement quantization, pruning, distillation, and compilation, and improve attention, KV caching, and memory layouts for multimodal models. A fast, impact-driven environment awaits.

Qualifications

  • 3+ years of experience in inference optimization, ML systems, or a closely related field.
  • Hands-on experience with modern ML stacks (PyTorch required; JAX a plus).
  • Strong understanding of compute, memory bandwidth, and I/O bottlenecks in large model inference.
  • Experience with optimization techniques: quantization, distillation, pruning, and compilation.
  • Familiarity with inference serving frameworks like Triton, TensorRT, vLLM, TorchServe.
  • Excellent debugging and measurement skills, with a track record of improvements.

Responsibilities

  • Own inference performance end-to-end: diagnose and improve latency, throughput, and efficiency of large foundation models in production.
  • Build systematic performance attribution: latency decomposition and bottleneck prioritization across model families.
  • Apply and develop optimization techniques including quantization, pruning, distillation, and model compilation.
  • Optimize attention mechanisms, KV caching, and memory layouts for multimodal models.
  • Collaborate with research engineers to translate innovations into deployment-ready implementations.
  • Develop benchmarking and regression detection infrastructure: latency baselines and throughput curves.
  • Work with kernel-level tooling (CUDA, Triton) to identify hotspots and tune custom kernels.

Skills

Inference optimization
ML systems
PyTorch
JAX
Quantization
Pruning
Distillation
Model compilation
CUDA
Triton
TensorRT
vLLM
TorchServe
Profiling
Latency optimization

Tools

TensorRT
CUDA
TorchServe
JAX
torch.compile

Job description

Rhoda AI is hiring an Inference Optimization MLE to build and operate systems that make foundation models run fast in production. Youll squeeze performance across cloud and on-robot targets, collaborating with research to close the training-deployment gap.

Youll own end-to-end optimization, implement quantization, pruning, distillation, and compilation, and improve attention, KV caching, and memory layouts for multimodal models. A fast, impact-driven environment awaits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Optimization Engineer for Real-World Robotics
Inference Optimization Engineer for Real-World Robotics

Rhoda AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Inference Optimization ML Engineer
Inference Optimization ML Engineer

Rhoda AI • Mountain View (WY)

On-site
USD 180,000 - 260,000
Senior Inference Optimization ML Engineer
Senior Inference Optimization ML Engineer

Rhoda AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Systems Engineer - Robot Learning Pipeline
Senior ML Systems Engineer - Robot Learning Pipeline

Socket.dev • Mountain View (CA)

On-site
USD 180,000 - 320,000
Real-Time AI Robotics Research Engineer
Real-Time AI Robotics Research Engineer

Rhoda AI • Mountain View (WY)

On-site
USD 150,000 - 230,000
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
ML Optimization Engineer: Training & Inference
ML Optimization Engineer: Training & Inference

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior DevOps Engineer - ML/Robotics Infra & Cloud
Senior DevOps Engineer - ML/Robotics Infra & Cloud

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior ML Training Systems Engineer, Large-Scale Robotics
Senior ML Training Systems Engineer, Large-Scale Robotics

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 200,000
Robotics AI Cloud Infrastructure Engineer
Robotics AI Cloud Infrastructure Engineer

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 190,000