Get more replies from employers
Send a job-specific resume in minutes.
Rhoda AI is hiring an Inference Optimization MLE to build and operate systems that make foundation models run fast in production. Youll squeeze performance across cloud and on-robot targets, collaborating with research to close the training-deployment gap.
Youll own end-to-end optimization, implement quantization, pruning, distillation, and compilation, and improve attention, KV caching, and memory layouts for multimodal models. A fast, impact-driven environment awaits.
Rhoda AI is hiring an Inference Optimization MLE to build and operate systems that make foundation models run fast in production. Youll squeeze performance across cloud and on-robot targets, collaborating with research to close the training-deployment gap.
Youll own end-to-end optimization, implement quantization, pruning, distillation, and compilation, and improve attention, KV caching, and memory layouts for multimodal models. A fast, impact-driven environment awaits.