Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
Relocation support SF/Shenzhen
Wellness benefits
Daily lunch and dinner in office
Unlimited compute budget
Codex and Claude credits
Travel

Job summary

Build AI, an in-person team in San Francisco, seeks a skilled ML/systems engineer to optimize inference performance, reduce latency and cost, and enable scaling of the data engine. You will work closely with research and product to ensure affordable, accurate models and efficient serving.

Ideal candidates think in dollars and tokens per second, are proficient in Python and C++/Rust, and are comfortable profiling and tuning GPUs/accelerators in a small, cost–constrained team.

Qualifications

  • Strong ML / systems engineer with real inference optimization experience (serving, compilers, CUDA/kernels, quantization, or similar).
  • Comfortable in Python and in C++ or Rust for performance‑critical paths.
  • You think in dollars and tokens/frames per second, not only in accuracy tables.
  • Familiar with PyTorch (or JAX) and profiling tools.
  • Able to work in a small research team shipping under cost pressure.

Responsibilities

  • Own inference performance: latency, throughput, and cost per unit of work (tokens, frames, or jobs).
  • Cut the 90% compute line: kernels, batching, quantization, compilation, serving, and hardware utilization.
  • Profile pipelines (Nsight, PyTorch Profiler, or equivalent), find the bottleneck, and ship the fix.
  • Collaborate with research and product so models are affordable to run at scale.
  • Build serving and eval path so experiments don’t inflate the inference bill.
  • Measure cost as a first‑class metric, not an afterthought when quality is done.

Skills

ML systems
Inference optimization
Python
C++/Rust
PyTorch
Profiling
Cost-aware thinking

Tools

Nsight
PyTorch Profiler
CUDA
Kernels

Job description

Build AI, an in-person team in San Francisco, seeks a skilled ML/systems engineer to optimize inference performance, reduce latency and cost, and enable scaling of the data engine. You will work closely with research and product to ensure affordable, accurate models and efficient serving.

Ideal candidates think in dollars and tokens per second, are proficient in Python and C++/Rust, and are comfortable profiling and tuning GPUs/accelerators in a small, cost–constrained team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Inference Systems Performance Engineer
Inference Systems Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Scale Low-Latency ML Inference Engineer (GPU/CUDA)
Scale Low-Latency ML Inference Engineer (GPU/CUDA)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Equity
Medical benefits (full coverage)
PTO & Hybrid work policy
+1
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Head of ML Systems & Inference
Head of ML Systems & Inference

Doist • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000