ML Inference & Systems Architect

Acceler8 Talent

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Acceler8 Talent in the Bay Area, CA is seeking a Member of Technical Staff - ML Systems & Inference to join onsite. You will help build the orchestration layer for next-generation AI workloads and work across ML systems, inference, and distributed systems teams.

The role focuses on producing production inference systems, optimizing scheduling and memory management, and improving KV cache efficiency to push AI infrastructure forward.

Qualifications

  • Strong software engineering fundamentals.
  • Experience with ML inference or model serving.
  • Knowledge of distributed systems and performance optimization.
  • Python and/or C++.
  • Experience with vLLM, TensorRT-LLM, CUDA, or similar is a plus.

Responsibilities

  • Build production inference systems, optimize scheduling and memory management.
  • Improve KV cache efficiency and performance of AI workloads.
  • Collaborate with compiler, kernel, and distributed systems engineers to advance AI infrastructure.

Skills

Strong software engineering
ML inference or model serving
Distributed systems and performance
Python
C++
vLLM
TensorRT-LLM
CUDA

Tools

CUDA
TensorRT-LLM
vLLM

Job description

Acceler8 Talent in the Bay Area, CA is seeking a Member of Technical Staff - ML Systems & Inference to join onsite. You will help build the orchestration layer for next-generation AI workloads and work across ML systems, inference, and distributed systems teams.

The role focuses on producing production inference systems, optimizing scheduling and memory management, and improving KV cache efficiency to push AI infrastructure forward.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Head of ML Systems & Inference
Head of ML Systems & Inference

Doist • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
ML Inference Systems Engineer
ML Inference Systems Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides
ML Systems Engineer — On-Site in Palo Alto, High-Impact
ML Systems Engineer — On-Site in Palo Alto, High-Impact

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
AI Inference Orchestration - Distributed Systems Engineer
AI Inference Orchestration - Distributed Systems Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000