Software Engineer, Distributed Training

River AI Inc.

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision insurance
Unlimited PTO
Relocation assistance

Job summary

River AI Inc. is seeking exceptional systems engineers to build the distributed training engines behind the River API, making fine-tuning and reinforcement learning fast, numerically correct, and reliable across large GPU clusters.

You will own training workload execution, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. You will collaborate with researchers and inference engineers to bring new learning methods into production and improve how

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
  • Hands-on experience building or substantially improving distributed model-training systems.
  • Strong proficiency in Python and a modern deep-learning framework, such as PyTorch or JAX.
  • Solid understanding of backpropagation, optimizers, mixed-precision training, and GPU memory management.
  • Strong debugging skills across concurrent execution, collective communication, and distributed failure recovery.
  • A highly collaborative mindset and a bias for action to push boundaries across the stack.

Responsibilities

  • Build and optimize distributed training for large dense and mixture-of-experts models, including low-rank adapter training.
  • Improve reinforcement-learning pipelines by coordinating sampling, reward computation, training updates, and weight transfer.
  • Optimize GPU memory use, parallelism, and communication to increase training throughput.
  • Implement reliable checkpointing, resumption, and worker recovery while preserving consistent training state.
  • Validate losses, gradients, and optimizer behavior, and diagnose numerical or distributed execution failures.
  • Partner with researchers to implement new algorithms and make them accessible through the River API.

Skills

Python
PyTorch
JAX
Distributed training
Debugging distributed systems

Education

Bachelor’s degree in Computer Science/Engineering

Tools

CUDA
C++
Rust
NCCL

Job description

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are looking for exceptional systems engineers to build the distributed training engines behind the River API. Your goal is to make fine-tuning and reinforcement learning fast, numerically correct, and reliable across large GPU clusters.

You will own the execution of training workloads, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. Working closely with researchers and inference engineers, you will bring new learning methods into production and improve how efficiently models use compute.

What You’ll Do
  • Build and optimize distributed training for large dense and mixture-of-experts models, including low-rank adapter training.
  • Improve reinforcement-learning pipelines by coordinating sampling, reward computation, training updates, and weight transfer.
  • Optimize GPU memory use, parallelism, and communication to increase training throughput.
  • Implement reliable checkpointing, resumption, and worker recovery while preserving consistent training state.
  • Validate losses, gradients, and optimizer behavior, and diagnose numerical or distributed execution failures.
  • Partner with researchers to implement new algorithms and make them accessible through the River API.
Skills & Qualifications

Minimum Qualifications:

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
  • Hands‑on experience building or substantially improving distributed model‑training systems.
  • Strong proficiency in Python and a modern deep‑learning framework, such as PyTorch or JAX.
  • Solid understanding of backpropagation, optimizers, mixed‑precision training, and GPU memory management.
  • Strong debugging skills across concurrent execution, collective communication, and distributed failure recovery.
  • A highly collaborative mindset and a bias for action to push boundaries across the stack.

Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

  • Experience with reinforcement‑learning infrastructure, rollout generation, or asynchronous training.
  • Familiarity with tensor, pipeline, expert, or data parallelism and their performance tradeoffs.
  • Work on LoRA, mixture‑of‑experts training, distributed optimizers, or activation checkpointing.
  • Experience with NCCL, communication profiling, and overlapping computation with data transfers.
  • Proficiency in C++, Rust, or CUDA, with experience investigating performance below the framework layer.
  • Contributions to training frameworks or a track record of operating large training runs.
  • Compensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Systems
Software Engineer, Inference Systems

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
+1
Distributed Training Systems Engineer
Distributed Training Systems Engineer

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Member of Program Staff, Data
Member of Program Staff, Data

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Research Engineer - Distributed Training
Research Engineer - Distributed Training

Prime Intellect • United States

Hybrid
USD 110,000 - 150,000
Competitive compensation including equity incentives
Flexible work arrangements
Visa sponsorship and relocation assistance
+2
Member of Technical Staff — Training
Member of Technical Staff — Training

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Member of Technical Staff – Model Training
Member of Technical Staff – Model Training

Inflection AI • Palo Alto (CA)

Hybrid
USD 175,000 - 350,000
Competitive stock options
Diverse medical, dental, and vision options
401k matching program
+3
Training Performance Engineer
Training Performance Engineer

Slope • San Francisco (CA)

On-site
USD 250,000 - 460,000
Relocation assistance
Flexible working hours
Collaborative work environment