Distributed Training Engineer — High-Perf GPU Scale

River AI

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
Relocation assistance
Visa sponsorship

Job summary

River AI in Palo Alto, CA, seeks exceptional systems engineers to build the distributed training engines behind the River API. Your work will focus on making fine-tuning fast, numerically correct, and reliable across large GPU clusters.

You will own execution of training workloads, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. Collaborate with researchers and inference engineers to deploy new learning methods and improve compute efficiency.

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Hands-on experience building or improving distributed model-training systems.
  • Strong proficiency in Python and a modern deep-learning framework such as PyTorch or JAX.
  • Solid understanding of backpropagation, optimizers, mixed-precision training, and GPU memory management.
  • Strong debugging skills across concurrent execution and distributed failure recovery.

Responsibilities

  • Build and optimize distributed training for large dense and mixture-of-experts models.
  • Improve reinforcement-learning pipelines by coordinating sampling, reward computation, training updates, and weight transfer.
  • Optimize GPU memory use, parallelism, and communication to increase training throughput.
  • Implement reliable checkpointing, resumption, and worker recovery while preserving training state.
  • Validate losses, gradients, and optimizer behavior, and diagnose numerical or distributed execution failures.
  • Partner with researchers to implement new algorithms and make them accessible through the River API.

Skills

Distributed training
Python
PyTorch
JAX
Backpropagation
GPU memory management
Distributed failure recovery

Education

Bachelors in CS/CE

Tools

NCCL
CUDA
C++
Rust

Job description

River AI in Palo Alto, CA, seeks exceptional systems engineers to build the distributed training engines behind the River API. Your work will focus on making fine-tuning fast, numerically correct, and reliable across large GPU clusters.

You will own execution of training workloads, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. Collaborate with researchers and inference engineers to deploy new learning methods and improve compute efficiency.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Distributed Training Systems Engineer
Distributed Training Systems Engineer

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Senior Systems Engineer for Scalable AI Training Infra
Senior Systems Engineer for Scalable AI Training Infra

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Systems Engineer - High-Performance AI Training
Systems Engineer - High-Performance AI Training

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Software Engineer, Distributed Training
Software Engineer, Distributed Training

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Relocation assistance
Visa sponsorship
Software Engineer, Distributed Training
Software Engineer, Distributed Training

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
GPU Kernel Engineer for Fast AI Training & Inference
GPU Kernel Engineer for Fast AI Training & Inference

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1
GPU Kernel Engineer — Accelerate AI Training & Inference
GPU Kernel Engineer — Accelerate AI Training & Inference

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
Software Engineer, River API
Software Engineer, River API

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3