Software Engineer, Inference Systems

River AI Inc.

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
Visa sponsorship

Job summary

River AI Inc. is seeking exceptional inference systems engineers to build engines that serve large models through the River API.

You will own the serving runtime, from request scheduling to distributed model execution, with a focus on latency, throughput, reliability, and cost. You will work with GPU kernel engineers, researchers, and infra engineers to bring models to production, optimize performance, and ensure model-version consistency across deployments.

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience building inference engines or performance-sensitive distributed services.
  • Strong understanding of transformer inference, GPU memory, concurrency, and networking.
  • Proficiency in Python and C++ or Rust.
  • Strong debugging and profiling skills across models, runtimes, and services.
  • A collaborative mindset and strong ownership of engineering outcomes.

Responsibilities

  • Optimize inference for dense and mixture-of-experts models, including fine-tuned models and adapters.
  • Improve batching, caching, and admission control to balance throughput, latency, memory use, and fairness.
  • Accelerate multi-GPU execution, communication, and model loading while preserving model-version consistency.
  • Improve RL sampling throughput while keeping samples and log probabilities tied to the correct model version.
  • Build reliable streaming, cancellation, and recovery under failures and overload.
  • Profile bottlenecks and validate improvements through reproducible performance and correctness tests.

Skills

Python
C++
Rust
Transformer inference
GPU memory
Profiling
Concurrency
Networking

Education

Bachelor’s degree in CS/CE or equivalent

Tools

CUDA
TensorRT-LLM
SGLang / vLLM

Job description

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are looking for exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory.

You will take ownership of the serving runtime, from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Your work will support both customer-facing inference and the sampling workloads that power reinforcement learning.

Working closely with GPU kernel engineers, researchers, and infrastructure engineers, you will bring new models into production and improve their performance across realistic workloads. You will measure success through latency, throughput, reliability, and cost, with careful attention to numerical correctness and model behavior.

What You’ll Do
  • Optimize inference for dense and mixture-of-experts models, including fine-tuned models and adapters.
  • Improve batching, caching, and admission control to balance throughput, latency, memory use, and fairness.
  • Accelerate multi-GPU execution, communication, and model loading while preserving model-version consistency.
  • Improve RL sampling throughput while keeping samples and log probabilities tied to the correct model version.
  • Build reliable streaming, cancellation, and recovery under failures and overload.
  • Profile bottlenecks and validate improvements through reproducible performance and correctness tests.
Skills & Qualifications

Minimum Qualifications:

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience building inference engines or performance-sensitive distributed services.
  • Strong understanding of transformer inference, GPU memory, concurrency, and networking.
  • Proficiency in Python and C++ or Rust.
  • Strong debugging and profiling skills across models, runtimes, and services.
  • A collaborative mindset and strong ownership of engineering outcomes.

Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

  • Experience extending SGLang, vLLM, TensorRT-LLM, or similar frameworks.
  • Work on advanced serving techniques, such as speculative decoding or disaggregated prefill and decode.
  • Familiarity with expert parallelism, GPU collectives, and quantized inference.
  • Experience with multi-adapter serving, dynamic checkpoint loading, or RL sampling.
  • Familiarity with CUDA graphs, custom kernels, and NVIDIA profiling tools.
  • Experience operating model-serving systems under production traffic.
  • Compensation: $200,000–$420,000 USD annual base pay, depending on experience and skills.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and support the process for the right candidate.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Distributed Training
Software Engineer, Distributed Training

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Inference Systems Engineer - Fast, Multi-GPU Model Serving
Inference Systems Engineer - Fast, Multi-GPU Model Serving

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Inference Performance Engineer
Inference Performance Engineer

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Staff + Senior Software Engineer, Inference
Staff + Senior Software Engineer, Inference

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits