ML Engineer, Inference Optimization

Build AI

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
Relocation support SF/Shenzhen
Wellness benefits
Daily lunch and dinner in office
Unlimited compute budget
Codex and Claude credits
Travel

Job summary

Build AI, an in-person team in San Francisco, seeks a skilled ML/systems engineer to optimize inference performance, reduce latency and cost, and enable scaling of the data engine. You will work closely with research and product to ensure affordable, accurate models and efficient serving.

Ideal candidates think in dollars and tokens per second, are proficient in Python and C++/Rust, and are comfortable profiling and tuning GPUs/accelerators in a small, cost–constrained team.

Qualifications

  • Strong ML / systems engineer with real inference optimization experience (serving, compilers, CUDA/kernels, quantization, or similar).
  • Comfortable in Python and in C++ or Rust for performance‑critical paths.
  • You think in dollars and tokens/frames per second, not only in accuracy tables.
  • Familiar with PyTorch (or JAX) and profiling tools.
  • Able to work in a small research team shipping under cost pressure.

Responsibilities

  • Own inference performance: latency, throughput, and cost per unit of work (tokens, frames, or jobs).
  • Cut the 90% compute line: kernels, batching, quantization, compilation, serving, and hardware utilization.
  • Profile pipelines (Nsight, PyTorch Profiler, or equivalent), find the bottleneck, and ship the fix.
  • Collaborate with research and product so models are affordable to run at scale.
  • Build serving and eval path so experiments don’t inflate the inference bill.
  • Measure cost as a first‑class metric, not an afterthought when quality is done.

Skills

ML systems
Inference optimization
Python
C++/Rust
PyTorch
Profiling
Cost-aware thinking

Tools

Nsight
PyTorch Profiler
CUDA
Kernels

Job description

About Build AI

Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.

Job Summary

Inference is about 90% of compute spend. Economics are heavily driven by inference optimization. We’re hiring someone to make inference cheaper, faster, and good enough that we can scale the data engine and the product without the GPU bill eating the company.

Key Responsibilities
  • Own inference performance: latency, throughput, and cost per unit of work (tokens, frames, or jobs)
  • Cut the 90% compute line: kernels, batching, quantization, compilation, serving, and hardware utilization
  • Profile pipelines (Nsight, PyTorch Profiler, or equivalent), find the real bottleneck, and ship the fix
  • Work with research and product so models that are accurate are also affordable to run at scale
  • Build the serving and eval path so experiments don’t hide the inference bill
  • Measure cost as a first‑class metric, not an afterthought once quality is “done”
You may be a good fit if you have (Must-have qualifications)
  • Strong ML / systems engineer with real inference optimization experience (serving, compilers, CUDA/kernels, quantization, or similar)
  • Comfortable in Python and in C++ or Rust for performance‑critical paths
  • You think in dollars and tokens/frames per second, not only in accuracy tables
  • Familiarity with PyTorch (or JAX) and with profiling tools
  • Comfortable in a small research team shipping under cost pressure
Strong candidates may also have experience with (Nice-to-have qualifications)
  • CUDA, kernels, compilers (TVM, MLIR, TensorRT), or quantization in production
  • You have owned GPU/accelerator cost as a first‑class metric
  • Serving stacks for video or large models
  • Understanding of memory hierarchy, data movement, and low‑precision compute
Benefits
  • Competitive pay
  • Medical, dental, and vision packages with generous premium coverage
  • $500 per month credit for waiving medical benefits
  • Housing subsidy of $2k per month for those living within walking distance of the office
  • Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)
  • Various wellness benefits covering fitness, mental health, and more
  • Daily lunch and dinner in our office
  • Unlimited compute budget subject to ROI justification
  • Unlimited Codex and Claude credits
  • Travel
How we're different

Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.

We are a fully in‑person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Build AI is an equal opportunity employer. We review every application.

Questions: research@build.ai

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Generalist
Software Engineer, Generalist

Worky • San Francisco (CA)

On-site
USD 135,000 - 210,000
Competitive pay
Health coverage
Medical credit for waiving benefits
+6
Member of Technical Staff
Member of Technical Staff

Build AI • San Francisco (CA)

On-site
USD 230,000 - 260,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy
+6
Software Engineer, Data Infrastructure & Pipelining
Software Engineer, Data Infrastructure & Pipelining

Build AI • San Francisco (CA)

On-site
USD 170,000 - 290,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy
+6
Software Engineer, Generalist
Software Engineer, Generalist

Build AI • San Francisco (CA)

On-site
USD 140,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy near office
+3
Software Engineer, Scaling
Software Engineer, Scaling

Build AI • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy for SF housing near HQ
+6
Data Scientist, Marketplace Incentives
Data Scientist, Marketplace Incentives

Build AI • San Francisco (CA)

On-site
USD 120,000 - 170,000
Competitive pay
Medical, dental, and vision coverage
Housing subsidy $2k per month near SF
+5
Lead Engineer, Data Platform
Lead Engineer, Data Platform

Build AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy for SF or Shenzhen
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Lead Engineer, ML Data Infrastructure & Systems
Lead Engineer, ML Data Infrastructure & Systems

Build AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy for SF office nearby
+6
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits