Research Member of Technical Staff- Efficient Modeling

Rhoda AI

Mountain View (CA)

On-site

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Rhoda AI is seeking a Research Scientist or Research Engineer in Mountain View, California, to enhance model efficiency for our intelligent robots. This role involves implementing model compression techniques and developing architectures for real-time applications. Ideal candidates are skilled in model compression and have experience with PyTorch.

Your contributions will bridge the gap between large-scale research models and real-time deployments, impacting the efficiency of every model trained and deployed.

Qualifications

  • Strong understanding of model compression and architectures for large models.
  • Hands-on experience with quantization and pruning applied to transformers.
  • Deep knowledge of efficiency gains in modern architectures.

Responsibilities

  • Research and implement model compression techniques.
  • Design efficient architectures for real-time inference on hardware.
  • Develop training strategies for better accuracy-efficiency tradeoffs.

Skills

Model compression
Efficient architectures
PyTorch
Quantization
Distillation

Education

PhD in ML, CS, or related field

Tools

TensorRT
CUDA

Job description

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $400M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

We're looking for a Research Scientist or Research Engineer focused on model efficiency — making our foundation world models faster, smaller, and more deployable without sacrificing capability. This work is critical to closing the gap between research-scale models and real-time operation on robot hardware.

What You’ll Do
  • Research and implement model compression techniques: quantization, pruning, structured sparsity, distillation, and low-rank approximation
  • Design efficient architectures and attention mechanisms suited to real-time inference on edge and robot hardware
  • Develop training strategies that produce better accuracy-efficiency tradeoffs from the start
  • Profile and benchmark models across hardware targets to identify and resolve efficiency bottlenecks
  • Build evaluation frameworks that measure capability retention after compression or architecture changes
  • Collaborate with training systems and deployment teams to ensure efficient models translate to faster real-world inference
  • Publish and present work at top-tier venues
What We’re Looking For
  • Strong understanding of model compression and efficient architectures for large models
  • Hands‑on experience with quantization, distillation, or pruning applied to transformers or large neural networks
  • Deep knowledge of where efficiency gains are possible in modern architectures
  • Proficiency with PyTorch and familiarity with hardware‑aware optimization (CUDA, TensorRT, or similar)
  • Ability to run principled experiments that characterize capability‑efficiency tradeoffs
Nice to Have (But Not Required)
  • PhD in ML, CS, or a related field — or equivalent research/engineering experience
  • Publication record at NeurIPS, ICML, ICLR, MLSys, or related venues
  • Experience with efficient video or multimodal model architectures
  • Familiarity with edge deployment targets (Jetson, custom ASICs, or mobile hardware)
  • Prior work on speculative decoding, early exit, or adaptive compute
  • Experience deploying compressed models on physical robots or latency-constrained systems
Why This Role
  • Bridge the gap between large‑scale research models and real‑time robot deployments
  • Your work determines whether frontier capabilities actually run on our hardware
  • High leverage: efficiency improvements benefit every model the team trains and deploys
  • Work at a rare intersection of deep learning research and systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Member of Technical Staff- Efficient Modeling
Research Member of Technical Staff- Efficient Modeling

Rhoda AI • Mountain View (WY)

On-site
USD 150,000 - 230,000
Senior Inference Optimization ML Engineer
Senior Inference Optimization ML Engineer

Rhoda AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 200,000
Research Member of Technical Staff- Deployment
Research Member of Technical Staff- Deployment

Rhoda AI • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 180,000
RESEARCHER, EFFICIENT INFERENCE
RESEARCHER, EFFICIENT INFERENCE

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 240,000
RESEARCHER, EFFICIENT INFERENCE
RESEARCHER, EFFICIENT INFERENCE

MLSys 2020 • San Francisco (CA)

On-site
USD 140,000 - 180,000
Research Member of Technical Staff- Post-training & Robot Learning
Research Member of Technical Staff- Post-training & Robot Learning

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Inference Optimization ML Engineer
Inference Optimization ML Engineer

Rhoda AI • Mountain View (WY)

On-site
USD 180,000 - 260,000