Senior ML Inference Systems Engineer

Acceler8 Talent

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Acceler8 Talent is seeking a Member of Technical Staff to join its AI infrastructure team in San Francisco. You will design and build production-grade ML inference and model serving systems, optimizing latency, throughput, and resource utilization across large-scale AI workloads.

You will work with inference runtimes such as vLLM or TensorRT-LLM, tackle memory management and scheduling challenges, and collaborate with compiler, kernel, and distributed systems engineers to push the performance of

Qualifications

  • Strong software engineering fundamentals.
  • Experience building ML inference or model serving systems in production.
  • Deep understanding of system performance and memory behaviour under production workloads.
  • Experience with inference runtimes such as vLLM, TensorRT-LLM, or custom serving frameworks.
  • Familiarity with batching, scheduling, concurrency, and KV cache management.
  • Experience profiling and optimising latency- and throughput-critical systems.
  • Strong Python and C++ development experience.

Responsibilities

  • Design and build production-grade ML inference and model serving systems
  • Optimise latency, throughput, and resource utilisation across large-scale AI workloads
  • Develop execution strategies around batching, scheduling, concurrency, and runtime optimisation
  • Improve KV cache management, memory efficiency, and model execution behaviour
  • Enable new model architectures and inference techniques to run efficiently in production
  • Partner closely with compiler, kernel, networking, and distributed systems engineers to drive end-to-end performance
  • Help shape the architecture of a next-generation AI inference platform powering production workloads at scale

Skills

Python
C++
ML inference
Model serving
Latency optimisation
Performance profiling
Concurrency
Batching
Scheduling
KV cache

Job description

Acceler8 Talent is seeking a Member of Technical Staff to join its AI infrastructure team in San Francisco. You will design and build production-grade ML inference and model serving systems, optimizing latency, throughput, and resource utilization across large-scale AI workloads.

You will work with inference runtimes such as vLLM or TensorRT-LLM, tackle memory management and scheduling challenges, and collaborate with compiler, kernel, and distributed systems engineers to push the performance of

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior ML Inference Engineer: High-Performance GPU Systems
Senior ML Inference Engineer: High-Performance GPU Systems

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000