Systems Engineer - High-Performance AI Training

River AI Inc.

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Unlimited PTO
Relocation assistance
Visa sponsorship

Job summary

River AI Inc. in Palo Alto is seeking exceptional systems engineers to build the high‑performance engines that train our models.

Your scope includes owning the core infrastructure stack, from writing custom GPU kernels to managing clusters of thousands of nodes, ensuring researchers stay focused on science, not bottlenecks. You will collaborate with researchers to rapidly implement, optimize, and scale experimental model architectures, delivering fast, reliable training and inference at scale

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
  • Deep expertise in systems-level languages (C, C++, or Rust) with a track record of writing performant, maintainable code.
  • Strong foundation in computer architecture, memory management, and concurrent programming.
  • Exceptional debugging skills, especially in distributed environments.
  • A highly collaborative mindset and a bias for action to push boundaries across the stack.

Responsibilities

  • Architect and deploy fault-tolerant distributed systems for training and inference workloads across clusters with thousands of nodes.
  • Design high-performance kernels to maximize tensor operation efficiency, memory throughput, and networking over InfiniBand/RDMA.
  • Profile systems end-to-end to resolve blockers across hardware, software, data loading pipelines, and collective communication primitives.
  • Partner directly with research scientists to rapidly implement, optimize, and scale experimental model architectures.

Skills

C/C++/Rust
Distributed systems
Performance optimization
Debugging distributed systems

Education

Bachelor’s degree in CS/CE or equivalent

Tools

GPU kernels
InfiniBand/RDMA
CUDA

Job description

River AI Inc. in Palo Alto is seeking exceptional systems engineers to build the high‑performance engines that train our models.

Your scope includes owning the core infrastructure stack, from writing custom GPU kernels to managing clusters of thousands of nodes, ensuring researchers stay focused on science, not bottlenecks. You will collaborate with researchers to rapidly implement, optimize, and scale experimental model architectures, delivering fast, reliable training and inference at scale

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Engineer for Scalable AI Training Infra
Senior Systems Engineer for Scalable AI Training Infra

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Distributed Training Engineer — High-Perf GPU Scale
Distributed Training Engineer — High-Perf GPU Scale

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Relocation assistance
Visa sponsorship
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Software Engineer, River API
Software Engineer, River API

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Distributed Training Systems Engineer
Distributed Training Systems Engineer

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Software Engineer, River API
Software Engineer, River API

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
GPU Kernel Engineer for Fast AI Training & Inference
GPU Kernel Engineer for Fast AI Training & Inference

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1
Software Engineer, Distributed Training
Software Engineer, Distributed Training

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Relocation assistance
Visa sponsorship
Software Engineer, Distributed Training
Software Engineer, Distributed Training

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
AI Engineer: Model Training, Inference & GPU Infra
AI Engineer: Model Training, Inference & GPU Infra

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 170,000 - 210,000