Research Member of Technical Staff - Training Platform

Rhoda AI

Mountain View (CA)

On-site

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Rhoda AI in Mountain View is seeking a Research Engineer to build and maintain a training platform that powers model development. The successful candidate will develop tooling for experiment management and optimize training processes across large-scale GPU clusters.

This role offers high visibility, direct feedback from researchers, and the chance to build systems that impact model training efficiency significantly. Candidates should have strong software engineering skills and familiarity with advanced training frameworks.

Qualifications

  • Strong software engineering skills with experience in MLOps or ML platform engineering.
  • Familiarity with distributed training frameworks like PyTorch DDP.
  • Experience building experiment tracking and artifact management systems.
  • Comfortable managing GPU cluster environments.
  • Strong reliability engineering instincts, including monitoring and alerting.

Responsibilities

  • Build and maintain training orchestration systems for large-scale distributed model training.
  • Develop experiment management tooling for job configuration and tracking.
  • Build observability infrastructure for monitoring training runs.
  • Optimize and automate the research iteration loop.
  • Manage job scheduling for efficient GPU compute utilization.
  • Collaborate with teams to support platform needs.

Skills

Software engineering skills
MLOps experience
Distributed training frameworks
Experiment tracking systems
GPU cluster management
Reliability engineering

Tools

Slurm
Kubernetes
AWS
GCP
Azure

Job description

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $400M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale‑up to make generalist robotics a reality.

We're looking for a Research Engineer to build and maintain the training platform that powers our model development — experiment orchestration, job management, observability, and the tooling that lets researchers move from idea to result as fast as possible.

What You'll Do
  • Build and maintain training orchestration systems for large-scale distributed model training across GPU clusters

  • Develop experiment management tooling: job configuration, tracking, reproducibility, and artifact management

  • Build observability infrastructure for training runs: loss curves, compute utilization, gradient statistics, and anomaly detection

  • Optimize and automate the research iteration loop from experiment launch to results analysis

  • Manage job scheduling and cluster utilization for efficient use of GPU compute

  • Build internal tooling and interfaces that help researchers move faster

  • Collaborate with training systems, data infrastructure, and research teams to support their platform needs

What We're Looking For
  • Strong software engineering skills with experience in MLOps or ML platform engineering

  • Familiarity with distributed training frameworks (PyTorch DDP, FSDP, DeepSpeed, Megatron, or similar)

  • Experience building experiment tracking, reproducibility, and artifact management systems

  • Comfortable managing and operating GPU cluster environments (Slurm, Kubernetes, or similar)

  • Strong reliability engineering instincts: monitoring, alerting, and failure recovery

Nice to Have (But Not Required)
  • Experience with training orchestration tools (Slurm, Ray, Kubernetes, or similar schedulers)

  • Familiarity with experiment tracking tools (Weights & Biases, MLflow, or custom solutions)

  • Experience supporting large model training pipelines (LLMs, VLMs, or video models)

  • Understanding of parallelism strategies and how they affect training efficiency and debugging

  • Experience with cloud-based training infrastructure (AWS, GCP, or Azure)

Why This Role
  • Your platform is the daily tool every researcher and engineer uses to train models

  • Improvements to training velocity and reliability compound across every experiment the team runs

  • High visibility with direct feedback from researchers and ML engineers

  • Build systems that scale from today's models to future frontier training runs

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 140,000 - 180,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 200,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Research Member of Technical Staff- Post-training & Robot Learning
Research Member of Technical Staff- Post-training & Robot Learning

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000
ML Platform Engineer – Training Orchestration
ML Platform Engineer – Training Orchestration

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 140,000
Research Member of Technical Staff- Data Infrastructure
Research Member of Technical Staff- Data Infrastructure

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior DevOps Engineer
Senior DevOps Engineer

Rhoda AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Research Member of Technical Staff- Robot Learning Systems & Reliability
Research Member of Technical Staff- Robot Learning Systems & Reliability

Rhoda • Mountain View (CA)

On-site
USD 180,000 - 240,000
Research Member of Technical Staff- Robot Learning Systems & Reliability
Research Member of Technical Staff- Robot Learning Systems & Reliability

Socket.dev • Mountain View (CA)

On-site
USD 180,000 - 320,000
Research Member of Technical Staff- Efficient Modeling
Research Member of Technical Staff- Efficient Modeling

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000