Inference Systems Engineer: Optimize AI Serving & Latency

Emploive

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Flexible work
Adaption Passport
Lunch Stipend
Well-Being

Job summary

Emploive is seeking an experienced ML systems engineer to own the cost and performance of our inference stack, shaping how we serve models as workloads, traffic, and hardware evolve. You will partner with the engineers operating the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level tuning.

You will implement profiling systems to show where time and memory are spent and tune routing between internal infrastructure and external providers using cost and capacity

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

Python
C++/Rust
CUDA
NCCL
Batching
ML systems
Performance engineering

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Emploive is seeking an experienced ML systems engineer to own the cost and performance of our inference stack, shaping how we serve models as workloads, traffic, and hardware evolve. You will partner with the engineers operating the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level tuning.

You will implement profiling systems to show where time and memory are spent and tune routing between internal infrastructure and external providers using cost and capacity

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Engineer: High-Throughput ML Serving
Inference Engineer: High-Throughput ML Serving

Adaption Labs, Inc. • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Flexible work: In-person collaboration
Adaption Passport
Lunch Stipend
+1
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
LLM Inference Performance Engineer: Speed & Efficiency
LLM Inference Performance Engineer: Speed & Efficiency

Baseten • United States

Remote
USD 150,000 - 210,000
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Inference Systems Engineer — Scalable, Reliable ML Serving
Inference Systems Engineer — Scalable, Reliable ML Serving

Speedrun Talent Network • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Low-Latency Inference Infrastructure Engineer
Low-Latency Inference Infrastructure Engineer

Elorian • Palo Alto (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,200 - 262,200
Health insurance
401(k) plan
Parental leave
+2
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000