Distributed Training Systems Engineer

River AI Inc.

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision insurance
Unlimited PTO
Relocation assistance

Job summary

River AI Inc. is seeking exceptional systems engineers to build the distributed training engines behind the River API, making fine-tuning and reinforcement learning fast, numerically correct, and reliable across large GPU clusters.

You will own training workload execution, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. You will collaborate with researchers and inference engineers to bring new learning methods into production and improve how

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
  • Hands-on experience building or substantially improving distributed model-training systems.
  • Strong proficiency in Python and a modern deep-learning framework, such as PyTorch or JAX.
  • Solid understanding of backpropagation, optimizers, mixed-precision training, and GPU memory management.
  • Strong debugging skills across concurrent execution, collective communication, and distributed failure recovery.
  • A highly collaborative mindset and a bias for action to push boundaries across the stack.

Responsibilities

  • Build and optimize distributed training for large dense and mixture-of-experts models, including low-rank adapter training.
  • Improve reinforcement-learning pipelines by coordinating sampling, reward computation, training updates, and weight transfer.
  • Optimize GPU memory use, parallelism, and communication to increase training throughput.
  • Implement reliable checkpointing, resumption, and worker recovery while preserving consistent training state.
  • Validate losses, gradients, and optimizer behavior, and diagnose numerical or distributed execution failures.
  • Partner with researchers to implement new algorithms and make them accessible through the River API.

Skills

Python
PyTorch
JAX
Distributed training
Debugging distributed systems

Education

Bachelor’s degree in Computer Science/Engineering

Tools

CUDA
C++
Rust
NCCL

Job description

River AI Inc. is seeking exceptional systems engineers to build the distributed training engines behind the River API, making fine-tuning and reinforcement learning fast, numerically correct, and reliable across large GPU clusters.

You will own training workload execution, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. You will collaborate with researchers and inference engineers to bring new learning methods into production and improve how

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Distributed Training
Software Engineer, Distributed Training

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Inference Systems Engineer - Fast, Multi-GPU Model Serving
Inference Systems Engineer - Fast, Multi-GPU Model Serving

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
+1
Software Engineer, Inference Systems
Software Engineer, Inference Systems

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
+1
GPU Kernel Engineer — Accelerate AI Training & Inference
GPU Kernel Engineer — Accelerate AI Training & Inference

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Principal AI Engineer Distributed Training Architect
Senior Principal AI Engineer Distributed Training Architect

Cerence Inc. • United States

Remote
USD 180,000 - 260,000
Senior AI Systems Architect: Scalable Distributed Training
Senior AI Systems Architect: Scalable Distributed Training

Cerence AI • United States

On-site
USD 180,000 - 240,000
Distributed AI Training Infrastructure Engineer
Distributed AI Training Infrastructure Engineer

Fireworks AI • United States

Remote
USD 130,000 - 210,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
LLM Distributed Training Engineer (1000+ GPUs)
LLM Distributed Training Engineer (1000+ GPUs)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000