Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent

San Francisco, Northern (CA, KY)

Hybrid

USD 170,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Acceler8 Talent is seeking a Member of Technical Staff to join our San Francisco onsite team building a next-gen AI inference platform. You will design and implement production-grade ML inference and model serving systems, optimizing latency and throughput for large-scale workloads.

You'll collaborate with compiler, kernel, networking, and distributed systems engineers to push performance and efficiency, supporting diverse hardware and heterogeneous compute in production environments.

Qualifications

  • Strong software engineering fundamentals.
  • Experience building ML inference or model serving systems in production.
  • Deep understanding of system performance, memory behaviour, and optimisation under production workloads.
  • Experience with inference runtimes such as vLLM, TensorRT-LLM, or custom serving frameworks.
  • Familiarity with batching, scheduling, concurrency, and KV cache management.
  • Experience profiling and optimising latency- and throughput-critical systems.
  • Strong Python and C++ development experience.
  • Comfortable working in a fast-moving, early-stage environment with significant ownership.

Responsibilities

  • Design and build production-grade ML inference and model serving systems
  • Optimise latency, throughput, and resource utilisation across large-scale AI workloads
  • Develop execution strategies around batching, scheduling, concurrency, and runtime optimisation
  • Improve KV cache management, memory efficiency, and model execution behaviour
  • Enable new model architectures and inference techniques to run efficiently in production
  • Partner closely with compiler, kernel, networking, and distributed systems engineers to drive end-to-end performance
  • Help shape the architecture of a next-generation AI inference platform powering production workloads at scale

Skills

Software engineering fundamentals
Production ML inference
Python
C++
Inference runtimes

Tools

vLLM
TensorRT-LLM
Custom serving frameworks

Job description

Acceler8 Talent is seeking a Member of Technical Staff to join our San Francisco onsite team building a next-gen AI inference platform. You will design and implement production-grade ML inference and model serving systems, optimizing latency and throughput for large-scale workloads.

You'll collaborate with compiler, kernel, networking, and distributed systems engineers to push performance and efficiency, supporting diverse hardware and heterogeneous compute in production environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference & Systems Architect
ML Inference & Systems Architect

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Staff Distributed Systems Engineer - AI Infrastructure
Staff Distributed Systems Engineer - AI Infrastructure

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
ML Systems Engineer: Scale Training & Inference
ML Systems Engineer: Scale Training & Inference

Doist • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive cash compensation
Startup equity
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Staff Engineer, Distributed AI Inference Systems (Equity)
Staff Engineer, Distributed AI Inference Systems (Equity)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Equity
Head of ML Systems & Inference
Head of ML Systems & Inference

Doist • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000