Founding Software Engineer (ML-Infra)

Reflection Robotics

Seattle (WA)

On-site

USD 140,000 - 200,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Reflection Robotics seeks a founding infrastructure engineer to create the base platform for training robotics foundation models. You will design a self-serve interface to launch training jobs across clouds, and own the orchestration that decides run destinations for efficiency and performance.

You will optimize GPU utilization, data pipelines, and checkpointing, while ensuring reliability and cost-efficiency. In-person Seattle work is required.

Qualifications

  • Strong systems engineering fundamentals used daily by engineers or robotics teams.
  • Experience with distributed training and large-scale ML workloads.
  • Proficiency with PyTorch.
  • Experience across AWS, GCP, Azure and cost/performance tradeoffs.
  • Comfort building and operating orchestration or scheduling systems.
  • Track record shipping infrastructure that holds up under daily use.
  • Ability to work in-person in Seattle on weekdays.

Responsibilities

  • Build a streamlined, self-serve way for robotics engineers to launch training jobs, abstracting away the underlying cloud provider or cluster
  • Design and build the orchestration layer that schedules and manages training jobs across multiple cloud providers
  • Continuously monitor and optimize training jobs for cost, GPU utilization, and throughput
  • Identify and eliminate bottlenecks in data loading, checkpointing, and distributed training that waste compute
  • Build tooling to make it easy to compare performance and cost across providers and hardware types
  • Set up monitoring and alerting so failed or underperforming jobs are caught quickly, not discovered days later
  • Work closely with the robotics engineering team to understand their workflows and remove friction wherever it shows up

Skills

Systems engineering
Distributed ML
PyTorch
Multi-cloud
Orchestration
Reliability
On-site Seattle
GPU Utilization
Cost optimization

Tools

Kubernetes

Job description

The Role

Training robot foundation models is expensive and iteration speed is everything: the faster our robotics engineers can launch a job, get results, and try the next idea, the faster the whole company moves. We need a founding engineer to build the training infrastructure that removes that gate, making it trivial to launch a training job, and making sure that job runs as efficiently as possible wherever it runs.

Concretely, that means building a system that lets robotics engineers launch training jobs with a single, simple interface, without needing to think about which cloud provider, which cluster, or which hardware is underneath it. You'll own the abstraction that decides where a job actually runs, whether that's optimizing for cost, availability, or performance across providers, and you'll be responsible for making sure GPUs aren't sitting idle and jobs aren't silently running slower than they should be.

This is a foundational role. The infrastructure you build will be the thing every robotics engineer touches every day, so the decisions you make about reliability, usability, and cost-efficiency will directly shape how fast this company can iterate on its core technology.

What You'll Do
  • Build a streamlined, self-serve way for robotics engineers to launch training jobs, abstracting away the underlying cloud provider or cluster

  • Design and build the orchestration layer that schedules and manages training jobs across multiple cloud providers

  • Continuously monitor and optimize training jobs for cost, GPU utilization, and throughput

  • Identify and eliminate bottlenecks in data loading, checkpointing, and distributed training that waste compute

  • Build tooling to make it easy to compare performance and cost across providers and hardware types

  • Set up monitoring and alerting so failed or underperforming jobs are caught quickly, not discovered days later

  • Work closely with the robotics engineering team to understand their workflows and remove friction wherever it shows up

What We're Looking For
Must-haves:
  • Strong systems engineering fundamentals, with experience building infrastructure that other engineers or robotics engineers depend on daily

  • Experience with distributed training and large-scale ML workloads

  • Proficiency with PyTorch

  • Experience working across multiple cloud providers (e.g. AWS, GCP, Azure) and reasoning about tradeoffs in cost and performance

  • Comfort building and operating orchestration or scheduling systems

  • A track record of shipping infrastructure that had to actually hold up under real, daily use

  • Ability to work in-person in Seattle on weekdays

Nice-to-haves:
  • Experience with distributed training frameworks (e.g. PyTorch Distributed, DeepSpeed, Ray)

  • Experience with Kubernetes or other container orchestration systems

  • Experience optimizing GPU utilization, data loading pipelines, or checkpointing at scale

  • Background in cloud cost optimization or FinOps for ML workloads

  • Experience building internal developer tools or platforms for robotics engineering teams

We care much more about how you think and what you've built than a specific degree or years-of-experience number. If you've done work that maps to this but doesn't check every box above, we'd still like to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer, Training Redwood City, CA Fulltime
ML Infrastructure Engineer, Training Redwood City, CA Fulltime

Dyna Robotics, Inc • Redwood City (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Founding ML Infra Engineer for Robotics Training Platform
Founding ML Infra Engineer for Robotics Training Platform

Reflection Robotics • Seattle (WA)

On-site
USD 140,000 - 200,000
Founding Robotics Engineer
Founding Robotics Engineer

Reflection Robotics • Seattle (WA)

On-site
USD 180,000 - 240,000
Founding Software Engineer [33397]
Founding Software Engineer [33397]

Stealth Startup • San Francisco (CA)

On-site
USD 200,000 - 280,000
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Member of Technical Staff - Training Platform
Research Member of Technical Staff - Training Platform

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 140,000
Systems Engineer
Systems Engineer

General Robotics • Redmond (WA)

On-site
USD 155,000 - 205,000
Staff AI Training Infrastructure Engineer
Staff AI Training Infrastructure Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 300,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)

Orbifold AI • Palo Alto (CA)

On-site
USD 170,000 - 230,000