Software Engineer, River API

River AI

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

36 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Unlimited PTO
Relocation assistance
Visa sponsorship

Job summary

River AI is seeking exceptional systems engineers to build the high-performance engines that train our models in Palo Alto, CA. You will own the core infrastructure stack, from writing custom GPU kernels to managing clusters of thousands of nodes, ensuring researchers can focus on science rather than system bottlenecks.

The role emphasizes fault-tolerant distributed systems, GPU kernel design, and collaboration with research scientists to scale experimental model architectures.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
  • Deep expertise in systems-level languages (C, C++, or Rust) with a track record of writing performant, maintainable code.
  • Strong foundation in computer architecture, memory management, and concurrent programming.
  • Exceptional debugging skills, especially when tackling complex, non-deterministic issues in distributed environments.
  • A highly collaborative mindset and a bias for action to push boundaries across the stack.

Responsibilities

  • Architect and deploy fault-tolerant distributed systems for training and inference workloads across clusters with thousands of nodes.
  • Design high-performance kernels to maximize tensor operation efficiency, memory throughput, and networking over InfiniBand/RDMA.
  • Profile systems end-to-end to resolve blockers across hardware, software, data loading pipelines, and collective communication primitives.
  • Partner directly with research scientists to rapidly implement, optimize, and scale experimental model architectures.

Skills

C/C++/Rust
Distributed systems
Computer architecture
Debugging
Collaborative mindset

Education

Bachelor's degree in CS/CE or equivalent

Tools

GPU kernels
InfiniBand/RDMA
PyTorch/JAX

Job description

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About The Role

We are looking for exceptional systems engineers to build the high-performance engines that train our models. Your goal is to make training at River fast, reliable, and massively scalable.

You will take ownership of our core infrastructure stack; from writing custom GPU kernels to managing clusters of thousands of nodes, ensuring our researchers can focus on science rather than system bottlenecks.

What You’ll Do
  • Architect and deploy fault-tolerant distributed systems for training and inference workloads across clusters with thousands of nodes.
  • Design high-performance kernels to maximize tensor operation efficiency, memory throughput, and networking over InfiniBand/RDMA.
  • Profile systems end-to-end to resolve blockers across hardware, software, data loading pipelines, and collective communication primitives.
  • Partner directly with research scientists to rapidly implement, optimize, and scale experimental model architectures.
Skills & Qualifications
Minimum Qualifications:
  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
  • Deep expertise in systems-level languages (C, C++, or Rust) with a track record of writing performant, maintainable code.
  • Strong foundation in computer architecture, memory management, and concurrent programming.
  • Exceptional debugging skills, especially when tackling complex, non-deterministic issues in distributed environments.
  • A highly collaborative mindset and a bias for action to push boundaries across the stack.
Preferred Qualifications:
  • Hands-on experience with modern AI frameworks (e.g., PyTorch, JAX) and tooling for large-scale model training.
  • Deep familiarity with modern GPU architectures (NVIDIA/AMD) and hardware constraints (HBM bandwidth, PCIe limits).
  • A proven track record of shipping and maintaining high-performance distributed systems or low-level software libraries.
  • Compensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Systems
Software Engineer, Inference Systems

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
+1
Software Engineer, Distributed Training
Software Engineer, Distributed Training

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1
Product Engineer, Personal AI
Product Engineer, Personal AI

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+2
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Performance Engineer, Hardware
Performance Engineer, Hardware

River AI Inc. • Palo Alto (CA), Austin (TX)

On-site
USD 200,000 - 420,000
Health benefits
Unlimited PTO
Relocation assistance
Member of Program Staff, Data
Member of Program Staff, Data

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
System Software Engineer - AI
System Software Engineer - AI

Delos Data • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits
Senior Systems Engineer for Scalable AI Training Infra
Senior Systems Engineer for Scalable AI Training Infra

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
RTL Design Engineer, Hardware
RTL Design Engineer, Hardware

River AI Inc. • Palo Alto (CA), Austin (TX)

On-site
USD 200,000 - 420,000
Health benefits
Dental benefits
Vision benefits
+2