Inference Systems Engineer - Fast, Multi-GPU Model Serving

River AI Inc.

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
Visa sponsorship

Job summary

River AI Inc. is seeking exceptional inference systems engineers to build engines that serve large models through the River API.

You will own the serving runtime, from request scheduling to distributed model execution, with a focus on latency, throughput, reliability, and cost. You will work with GPU kernel engineers, researchers, and infra engineers to bring models to production, optimize performance, and ensure model-version consistency across deployments.

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience building inference engines or performance-sensitive distributed services.
  • Strong understanding of transformer inference, GPU memory, concurrency, and networking.
  • Proficiency in Python and C++ or Rust.
  • Strong debugging and profiling skills across models, runtimes, and services.
  • A collaborative mindset and strong ownership of engineering outcomes.

Responsibilities

  • Optimize inference for dense and mixture-of-experts models, including fine-tuned models and adapters.
  • Improve batching, caching, and admission control to balance throughput, latency, memory use, and fairness.
  • Accelerate multi-GPU execution, communication, and model loading while preserving model-version consistency.
  • Improve RL sampling throughput while keeping samples and log probabilities tied to the correct model version.
  • Build reliable streaming, cancellation, and recovery under failures and overload.
  • Profile bottlenecks and validate improvements through reproducible performance and correctness tests.

Skills

Python
C++
Rust
Transformer inference
GPU memory
Profiling
Concurrency
Networking

Education

Bachelor’s degree in CS/CE or equivalent

Tools

CUDA
TensorRT-LLM
SGLang / vLLM

Job description

River AI Inc. is seeking exceptional inference systems engineers to build engines that serve large models through the River API.

You will own the serving runtime, from request scheduling to distributed model execution, with a focus on latency, throughput, reliability, and cost. You will work with GPU kernel engineers, researchers, and infra engineers to bring models to production, optimize performance, and ensure model-version consistency across deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Systems
Software Engineer, Inference Systems

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
+1
GPU Kernel Engineer — Accelerate AI Training & Inference
GPU Kernel Engineer — Accelerate AI Training & Inference

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Distributed Training Systems Engineer
Distributed Training Systems Engineer

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision insurance
Unlimited PTO
Relocation assistance
Inference Infra Engineer: Scale Low-Latency AI Serving
Inference Infra Engineer: Scale Low-Latency AI Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Senior Model Serving Engineer – Remote AI Infra
Senior Model Serving Engineer – Remote AI Infra

United States Digital Space LLC • United States

Remote
USD 74,000 - 98,000
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Remote Model Serving Engineer - Scale ML Inference
Remote Model Serving Engineer - Scale ML Inference

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 74,000 - 98,000
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000