Senior Systems Engineer for Scalable AI Training Infra

River AI

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

38 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Unlimited PTO
Relocation assistance
Visa sponsorship

Job summary

River AI is seeking exceptional systems engineers to build the high-performance engines that train our models in Palo Alto, CA. You will own the core infrastructure stack, from writing custom GPU kernels to managing clusters of thousands of nodes, ensuring researchers can focus on science rather than system bottlenecks.

The role emphasizes fault-tolerant distributed systems, GPU kernel design, and collaboration with research scientists to scale experimental model architectures.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
  • Deep expertise in systems-level languages (C, C++, or Rust) with a track record of writing performant, maintainable code.
  • Strong foundation in computer architecture, memory management, and concurrent programming.
  • Exceptional debugging skills, especially when tackling complex, non-deterministic issues in distributed environments.
  • A highly collaborative mindset and a bias for action to push boundaries across the stack.

Responsibilities

  • Architect and deploy fault-tolerant distributed systems for training and inference workloads across clusters with thousands of nodes.
  • Design high-performance kernels to maximize tensor operation efficiency, memory throughput, and networking over InfiniBand/RDMA.
  • Profile systems end-to-end to resolve blockers across hardware, software, data loading pipelines, and collective communication primitives.
  • Partner directly with research scientists to rapidly implement, optimize, and scale experimental model architectures.

Skills

C/C++/Rust
Distributed systems
Computer architecture
Debugging
Collaborative mindset

Education

Bachelor's degree in CS/CE or equivalent

Tools

GPU kernels
InfiniBand/RDMA
PyTorch/JAX

Job description

River AI is seeking exceptional systems engineers to build the high-performance engines that train our models in Palo Alto, CA. You will own the core infrastructure stack, from writing custom GPU kernels to managing clusters of thousands of nodes, ensuring researchers can focus on science rather than system bottlenecks.

The role emphasizes fault-tolerant distributed systems, GPU kernel design, and collaboration with research scientists to scale experimental model architectures.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Training Systems Engineer
Distributed Training Systems Engineer

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Software Engineer, River API
Software Engineer, River API

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Software Engineer, Distributed Training
Software Engineer, Distributed Training

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
Inference Systems Engineer - Fast, Multi-GPU Model Serving
Inference Systems Engineer - Fast, Multi-GPU Model Serving

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
+1
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Infra Architect: Scalable GPU & Edge Platforms
Senior AI Infra Architect: Scalable GPU & Edge Platforms

Seekr • San Francisco (CA)

Hybrid
USD 190,000 - 260,000
Equity ownership – RSUs
Unlimited PTO
Flexible hybrid work environment
+3
RL Infrastructure Engineer for Scalable GPU Training
RL Infrastructure Engineer for Scalable GPU Training

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000