Inference Systems Engineer — High-Performance AI Serving

River AI

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Unlimited PTO
Relocation assistance
Visa sponsorship

Job summary

River AI in Palo Alto is seeking exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory.

You will own the serving runtime from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Working closely with GPU kernel engineers, researchers, and infrastructure engineers, you will

Qualifications

  • Bachelor’s degree in CS, CE, or equivalent practical experience.
  • Experience building inference engines or performance-sensitive distributed services.
  • Strong knowledge of transformer inference, GPU memory, concurrency, and networking.
  • Proficiency in Python and C++ or Rust.
  • Strong debugging and profiling skills across models, runtimes, and services.
  • Collaborative mindset with ownership of outcomes.

Responsibilities

  • Optimize inference for dense and mixture-of-experts models.
  • Improve batching, caching, and admission control for throughput and latency.
  • Accelerate multi-GPU execution while preserving model-version consistency.
  • Improve RL sampling throughput with correct model versioning.
  • Build reliable streaming, cancellation, and recovery under load.
  • Profile bottlenecks and validate improvements with reproducible tests.

Skills

Python
C++
Rust
Transformer inference
GPU memory
Concurrency
Networking
Debugging
Profiling
Distributed services

Education

Bachelor's degree in CS/CE

Tools

SGLang
vLLM
TensorRT-LLM
CUDA
NVIDIA profiling tools

Job description

River AI in Palo Alto is seeking exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory.

You will own the serving runtime from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Working closely with GPU kernel engineers, researchers, and infrastructure engineers, you will

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Systems Engineer - High-Performance AI Training
Systems Engineer - High-Performance AI Training

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Software Engineer, Inference Systems
Software Engineer, Inference Systems

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Systems Engineer for Scalable AI Training Infra
Senior Systems Engineer for Scalable AI Training Infra

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Distributed Training Engineer — High-Perf GPU Scale
Distributed Training Engineer — High-Perf GPU Scale

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Relocation assistance
Visa sponsorship
Inference Infra Engineer: Scale Low-Latency AI Serving
Inference Infra Engineer: Scale Low-Latency AI Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Inference Systems Engineer, High-Throughput AI Serving
Inference Systems Engineer, High-Throughput AI Serving

Future Ventures • Palo Alto (CA)

On-site
USD 135,000 - 160,000
Comprehensive medical, vision, dental coverage
401(k) retirement plan
Paid parental leave
+1
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Remote Model Serving Engineer: Scalable AI Inference
Remote Model Serving Engineer: Scalable AI Inference

Socket.dev • Ann Arbor (MI)

On-site
USD 74,000 - 98,000
GPU Kernel Engineer — Accelerate AI Training & Inference
GPU Kernel Engineer — Accelerate AI Training & Inference

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
GPU Kernel Engineer for Fast AI Training & Inference
GPU Kernel Engineer for Fast AI Training & Inference

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1