Research Engineer

Dyna Robotics

Redwood City (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Dyna Robotics is seeking a Research Engineer to architect and build end-to-end training infrastructure for a multi-cloud GPU fleet. You will ensure GPUs are utilized efficiently, runs are reproducible, and researchers can iterate with a single command.

You will own training infrastructure end-to-end, working on scalable systems for large multimodal models and real-time robot data processing, with emphasis on performance, reliability, and deployable inference.

Qualifications

  • 7+ years of engineering experience in HPC or ML infrastructure.
  • Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate).
  • Hands-on cloud GPU management (GCP/AWS) and container orchestration (Kubernetes).
  • Strong understanding of distributed systems, NCCL inter-node communication, memory management.

Responsibilities

  • Scale distributed training infrastructure for large-scale GPU clusters, including sharding, checkpointing, and memory optimization.
  • Build a research codebase and job scheduling system to enable fast iteration and reliable retries.
  • Create high-throughput data pipelines for terabytes of multimodal robot data, ensuring GPUs are never starved.
  • Develop low-latency production inference pipelines with quantization and model compilation (TensorRT, Triton).
  • Profile GPU utilization and memory to optimize performance across the compute fleet.

Job description

Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model with top-in-industry generalization and real-world performance. Already deployed with customers across multiple industries, our robots do commercial-grade work in the physical world. Our team comes from Google DeepMind, Meta, and Cruise, and we're backed by CRV, First Round, and other leading investors.

The Role

As a Research Engineer, you will architect and build the systems that turn our multi-cloud GPU fleet into a training engine our researchers love. Your charter is singular and broad: own training infrastructure end-to-end so that every GPU is busy, every run is reproducible, and every researcher's next experiment is one command away.

What You’ll Do
  • Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters. You’ll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.

  • Optimize Researcher Ergonomics: Build a research codebase and job scheduling system (Kubernetes/SLURM) that prioritizes fast iteration, automated retries, and seamless failure recovery.

  • High-Performance Data Handling: Design high-throughput pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals), ensuring dataloaders never starve the GPUs.

  • Production Inference: Build low-latency inference pipelines for real-time robot control. You’ll apply quantization, distillation, and model compilation (TensorRT, Triton) to move models from the lab to the physical world.

  • Deep Systems Profiling: Dive into the weeds of GPU utilization, I/O bottlenecks, and memory fragmentation to squeeze every bit of performance out of our expanding compute fleet.

What You’ll Bring
  • 7+ Years of Engineering: With a track record of leading technical projects in high-performance computing (HPC) or ML infrastructure.

  • ML Systems Mastery: Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate). You understand the nuances of mixed precision and gradient accumulation.

  • Infrastructure Expertise: Hands-on experience managing cloud GPU environments (GCP/AWS) and container orchestration (Kubernetes).

  • Low-Level Intuition: A fundamental understanding of distributed systems, including race conditions, memory management, and NCCL/inter-node communication.

  • Ownership Mindset: You don't just "deploy" code; you design, build, and operate systems end-to-end to unblock fast-moving research.

Bonus Points For
  • Experience with Robotics Data Formats (MCAP, Protobuf) or multimodal models (VLAs).

  • Deep ML systems experience: custom kernels (Triton), compilers, or runtime optimization.

  • Experience as a founding or early-stage infrastructure hire.

At Dyna Robotics, we build technology for the real world, which requires a team as diverse as the environments our robots inhabit. We are an equal opportunity employer committed to technical rigor and mutual respect.

Don’t let a checklist stop you. Data shows that underrepresented groups often only apply if they meet 100% of the criteria. We value problem-solving and grit over keyword matching. If you’re passionate about robotics, no matter your discipline, we want to hear from you, even if you don't check every box.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Applications Redwood City, CA Fulltime
Software Engineer, Applications Redwood City, CA Fulltime

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

On-site
USD 150,000 - 230,000
Applied Researcher - Deployment Intelligence & Continuous Learning
Applied Researcher - Deployment Intelligence & Continuous Learning

Dyna Robotics • Redwood City (CA)

On-site
USD 140,000 - 210,000
Research Engineer/Scientist, Simulation
Research Engineer/Scientist, Simulation

Dyna Robotics • Redwood City (CA)

On-site
USD 180,000 - 240,000
Software Engineer, Data Infra
Software Engineer, Data Infra

Dyna Robotics • Redwood City (CA)

On-site
USD 120,000 - 160,000
Research Engineer/Scientist, Simulation Redwood City, CA Fulltime
Research Engineer/Scientist, Simulation Redwood City, CA Fulltime

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Head of Engineering
Head of Engineering

Dyna Robotics • Redwood City (CA)

On-site
USD 130,000 - 170,000
Exceptional Software Engineer Redwood City, CA Fulltime
Exceptional Software Engineer Redwood City, CA Fulltime

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Robotics Hardware Reliability Engineer Redwood City, CA Fulltime
Robotics Hardware Reliability Engineer Redwood City, CA Fulltime

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Robotics Hardware Engineer, Reliability & Bring-up
Robotics Hardware Engineer, Reliability & Bring-up

Dyna Robotics • Redwood City (CA)

On-site
USD 140,000 - 180,000
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000