Senior ML Inference Engineer: High-Performance GPU Systems

Acceler8 Talent

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Acceler8 Talent is recruiting an ML Inference Engineer for a Stanford-spun AI startup in San Francisco that is building an eight-figure revenue and growth trajectory. You will design, implement, and optimize the infrastructure powering large-scale LLM workloads and real-time model serving.

Ideal candidates combine Python/C++ proficiency with distributed systems experience, PyTorch expertise, and a passion for low-latency, GPU-accelerated inference at production scale.

Qualifications

  • Experience with high-performance computing or distributed systems.
  • Strong knowledge of Python and C++ for performance-sensitive stack.
  • Experience with LLM inference and model serving.
  • Experience optimizing GPU-heavy workloads and real-time serving.

Responsibilities

  • Build the infrastructure that serves large-scale LLM workloads.
  • Push the limits of latency, throughput and GPU efficiency.
  • Design distributed inference across single and multi-GPU systems.
  • Improve GPU scheduling, orchestration and resource utilisation.
  • Profile and remove bottlenecks across compute, memory and networking.
  • Scale production workloads across Kubernetes and GPU clusters.
  • Make low-level architecture decisions where milliseconds matter.

Skills

Python
C++
Distributed systems
LLM inference
PyTorch

Tools

CUDA
NCCL
Triton
vLLM
TensorRT-LLM
SGLang

Job description

Acceler8 Talent is recruiting an ML Inference Engineer for a Stanford-spun AI startup in San Francisco that is building an eight-figure revenue and growth trajectory. You will design, implement, and optimize the infrastructure powering large-scale LLM workloads and real-time model serving.

Ideal candidates combine Python/C++ proficiency with distributed systems experience, PyTorch expertise, and a passion for low-latency, GPU-accelerated inference at production scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Systems Engineer
Senior ML Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Machine Learning Inference Engineer
Machine Learning Inference Engineer

Oscar Technology • San Francisco (CA)

Hybrid
USD 225,000 - 275,000
Equity
401k matching
Medical coverage
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior ML Performance Engineer — Ultra-Fast Inference + Equity
Senior ML Performance Engineer — Ultra-Fast Inference + Equity

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
High-Performance ML Inference Engineer
High-Performance ML Inference Engineer

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2