Machine Learning Engineer (Inference)

Acceler8 Talent

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Acceler8 Talent is recruiting an ML Inference Engineer for a Stanford-spun AI startup in San Francisco that is building an eight-figure revenue and growth trajectory. You will design, implement, and optimize the infrastructure powering large-scale LLM workloads and real-time model serving.

Ideal candidates combine Python/C++ proficiency with distributed systems experience, PyTorch expertise, and a passion for low-latency, GPU-accelerated inference at production scale.

Qualifications

  • Experience with high-performance computing or distributed systems.
  • Strong knowledge of Python and C++ for performance-sensitive stack.
  • Experience with LLM inference and model serving.
  • Experience optimizing GPU-heavy workloads and real-time serving.

Responsibilities

  • Build the infrastructure that serves large-scale LLM workloads.
  • Push the limits of latency, throughput and GPU efficiency.
  • Design distributed inference across single and multi-GPU systems.
  • Improve GPU scheduling, orchestration and resource utilisation.
  • Profile and remove bottlenecks across compute, memory and networking.
  • Scale production workloads across Kubernetes and GPU clusters.
  • Make low-level architecture decisions where milliseconds matter.

Skills

Python
C++
Distributed systems
LLM inference
PyTorch

Tools

CUDA
NCCL
Triton
vLLM
TensorRT-LLM
SGLang

Job description

ML Inference Engineer
Build the infrastructure that makes cutting-edge AI fast enough to work at scale.

We’re hiring anML Inference Engineerfor a Stanford-spun AI startup in San Francisco that has already grown to8-figure revenue.

The team is rebuilding itsLLM inference stack from the ground up, solving challenging systems problems around GPU performance, distributed compute and real-time model serving.

This role is for engineers who enjoy going deep onperformance, infrastructure, and optimisation.

The role
  • Build the infrastructure that serveslarge-scale LLM workloads
  • Push the limits oflatency, throughput and GPU efficiency
  • Design distributed inference acrosssingle and multi-GPU systems
  • Improve GPU scheduling, orchestration and resource utilisation
  • Profile and remove bottlenecks acrosscompute, memory and networking
  • Scale production workloads acrossKubernetes and GPU clusters
  • Make low-level architecture decisions wheremilliseconds matter
What we're looking for
  • StrongPython and/or C++
  • Experience withdistributed systems or high-performance computing
  • Knowledge ofLLM inference and model serving
  • Experience optimising GPU-heavy workloads
  • Exposure toCUDA, NCCL or Triton
  • Strong understanding ofPyTorchand modern ML infrastructure
  • Experience withvLLM, TensorRT-LLM, SGLang or similar
  • Knowledge of techniques such asquantisation, batching, KV caching and parallelism
Why join?
  • Tackle genuinely difficultAI infrastructure problems
  • Work on systems operating atreal production scale
  • Join a fast-growing company already at8-figure revenue
  • Significant ownership over anew inference architecture
  • Work at the intersection ofLLMs, GPUs and distributed systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior ML Inference Engineer: High-Performance GPU Systems
Senior ML Inference Engineer: High-Performance GPU Systems

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 180,000 - 240,000
Health & Wellbeing
Personal & Professional Development
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage