Software Engineer - ML Infrastructure

Watney

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Watney Robotics is hiring a ML Infrastructure software engineer to architect high-performance systems for large-scale multimodal training. You’ll own training/inference infrastructure, scale distributed training across heterogeneous GPU/TPU clusters, and optimize memory, throughput, and resource costs.

This role sits at the intersection of deep learning, hardware acceleration, and scalable infrastructure, enabling rapid experiments and robust production pipelines.

Qualifications

  • Bring a experience building machine learning platforms and large-scale distributed training.
  • Possess deep professional experience with distributed training backbones (FSDP, DeepSpeed, Megatron, Ray Train) or large-scale inference serving layers (vLLM, Triton, Ray Serve).
  • Exhibit fluency in Python alongside Rust or C/C++, with a strong mathematical background and practical knowledge of GPU kernel optimization or network topologies.
  • Have experience navigating structural edge-case hardware bottlenecks, specifically regarding video decoding, multimodal alignment, or high-throughput real-time playback.

Responsibilities

  • <p><b>Own Training & Inference Infrastructure:</b> Design and maintain multi-tenant scheduling systems that automatically place training and inference jobs based on hardware topology, cost, and priority, while enforcing fair resource sharing and preemption policies.</p>
  • <p><b>Scale Distributed Training:</b> Partner with researchers to scale JAX and PyTorch-based training loops across heterogeneous GPU/TPU clusters with minimal friction, ensuring rock-solid checkpointing and metrics collection.</p>
  • <p><b>Optimize Performance & Hardware Bounds:</b> Profile and improve memory usage, device utilization, throughput, and distributed synchronization, specifically navigating edge hardware bottlenecks like on-chip video decoders and memory bandwidth.</p>
  • <p><b>Enable Rapid Iteration:</b> Build clean abstractions for launching, monitoring, debugging, and reproducing experiments so researchers can submit massive jobs without needing to manage underlying cluster state.</p>
  • <p><b>Contribute to Core Training Code:</b> Evolve our core JAX model code and training pipelines to natively support new architectures, multimodal video/telemetry data streams, and robust evaluation metrics.</p>
  • <p><b>Manage Compute Resources:</b> Ensure highly efficient allocation and utilization of massive cloud-based compute clusters while aggressively monitoring and controlling resource costs.</p>

Skills

Distributed training
Python
Rust/C++
GPU optimization
Cluster management

Tools

DeepSpeed
Megatron
Ray Train
Triton
vLLM

Job description

Our Mission

Expand human ambition in the physical world.


Critical infrastructure is constrained by labor shortages, hazardous working conditions, and operational complexity. Watney builds and deploys autonomous robotic systems that increase the speed and capacity of buildout, starting with data centers.


About the Role

At Watney, ML Infrastructure software engineers build the high-performance foundations that allow our perception and intelligence models to scale. You will architect the high-performance computing foundation that powers our physical intelligence models.


In this role, you’ll own the infrastructure required for large-scale multimodal training, which includes cluster orchestration, optimizing JAX-based pipelines that must ingest and stream video data, and transforming experimental architectures into reliable, highly distributed production training runs.


This is a high-leverage systems role at the intersection of deep learning, advanced hardware acceleration, and scalable cluster infrastructure.


What You’ll Do



  • Own Training & Inference Infrastructure: Design and maintain multi-tenant scheduling systems that automatically place training and inference jobs based on hardware topology, cost, and priority, while enforcing fair resource sharing and preemption policies.




  • Scale Distributed Training: Partner with researchers to scale JAX and PyTorch-based training loops across heterogeneous GPU/TPU clusters with minimal friction, ensuring rock-solid checkpointing and metrics collection.




  • Optimize Performance & Hardware Bounds: Profile and improve memory usage, device utilization, throughput, and distributed synchronization, specifically navigating edge hardware bottlenecks like on-chip video decoders and memory bandwidth.




  • Enable Rapid Iteration: Build clean abstractions for launching, monitoring, debugging, and reproducing experiments so researchers can submit massive jobs without needing to manage underlying cluster state.




  • Contribute to Core Training Code: Evolve our core JAX model code and training pipelines to natively support new architectures, multimodal video/telemetry data streams, and robust evaluation metrics.




  • Manage Compute Resources: Ensure highly efficient allocation and utilization of massive cloud-based compute clusters while aggressively monitoring and controlling resource costs.




You May Be a Good Fit If You:



  • Bring a experience building machine learning platforms and large-scale distributed training




  • Possess deep professional experience with distributed training backbones (FSDP, DeepSpeed, Megatron, Ray Train) or large-scale inference serving layers (vLLM, Triton, Ray Serve).




  • Exhibit fluency in Python alongside Rust or C/C++, with a strong mathematical background and practical knowledge of GPU kernel optimization or network topologies.




  • Have experience navigating structural edge-case hardware bottlenecks, specifically regarding video decoding, multimodal alignment, or high-throughput real-time playback.




We’re committed to building a diverse, inclusive team. At Watney Robotics, we welcome people of all backgrounds and identities, and we make hiring decisions based on skills, experience, and potential. If you’re passionate about robotics but don’t meet every requirement, we still encourage you to apply!


Curious to learn more?

Follow us here on X and LinkedIn

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer - ML Infrastructure
Staff Software Engineer - ML Infrastructure

Watney • San Francisco (CA)

On-site
USD 130,000 - 190,000
Staff Software Engineer - ML Infrastructure
Staff Software Engineer - ML Infrastructure

Watney Robotics Inc • San Francisco (CA)

On-site
USD 140,000 - 210,000
Software Engineer
Software Engineer

Watney • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer
Software Engineer

Watney Robotics Inc • San Francisco (CA)

On-site
USD 120,000 - 180,000
Software Engineer - Perception
Software Engineer - Perception

Watney Robotics Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer - Perception
Software Engineer - Perception

Watney • San Francisco (CA)

On-site
USD 150,000 - 210,000
Staff Embedded Engineer
Staff Embedded Engineer

Watney Robotics Inc • San Francisco (CA)

On-site
USD 120,000 - 180,000
Staff Software Engineer - Full Stack
Staff Software Engineer - Full Stack

Watney Robotics Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff ML Infra Engineer: Scalable Training & Pipelines
Staff ML Infra Engineer: Scalable Training & Pipelines

Watney Robotics Inc • San Francisco (CA)

On-site
USD 140,000 - 210,000
Staff ML Infra Engineer: Scalable Training & Pipelines
Staff ML Infra Engineer: Scalable Training & Pipelines

Watney • San Francisco (CA)

On-site
USD 130,000 - 190,000