AI Inference Engineer

Acceler8 Talent

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

42 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Acceler8 Talent in San Francisco seeks an ML Inference Engineer to redesign the infrastructure behind its production AI platform. You will architect high-performance systems for serving LLMs, optimize inference for latency and throughput, and scale GPU workloads across a Kubernetes-based stack.

You will work close to the hardware and software stack, balancing trade-offs for impactful improvements at scale.

Qualifications

  • Strong background in ML systems, distributed computing, HPC or performance engineering.
  • Experience working close to infrastructure that runs modern ML models.
  • Proficiency in Python and/or C++.
  • Solid understanding of PyTorch and GPU computing.
  • Experience with CUDA, NCCL or Triton.
  • Hands-on work with inference frameworks such as vLLM, TensorRT-LLM or SGLang.

Responsibilities

  • Architect high-performance systems for serving LLMs at production scale.
  • Optimize inference for latency, throughput and cost efficiency.
  • Build distributed execution across multi-GPU and multi-node environments.
  • Improve GPU resource scheduling, allocation and utilization.
  • Profile the inference stack to identify bottlenecks across compute, memory and networking.
  • Scale GPU workloads across Kubernetes-based infrastructure.
  • Optimize techniques including batching, quantisation, KV caching and parallelism.

Skills

ML systems
Distributed computing
HPC
Performance engineering
Python
C++
PyTorch
GPU computing
CUDA
NCCL
Triton

Tools

CUDA Toolkit
NCCL
Triton

Job description

A Stanford-spun AI company in San Francisco is hiring an ML Inference Engineer to help redesign the infrastructure behind its production AI platform.

What you'll be solving

  • Architect high-performance systems for serving LLMs at production scale
  • Optimise inference for latency, throughput and cost efficiency
  • Build distributed execution across multi-GPU and multi-node environments
  • Improve how GPU resources are scheduled, allocated and utilised
  • Profile the inference stack to identify bottlenecks across compute, memory and networking
  • Scale GPU workloads across Kubernetes-based infrastructure
  • Optimise techniques including batching, quantisation, KV caching and parallelism
  • Make systems-level trade-offs where small improvements can have a major impact at scale

What you'll bring

  • You'll have a strong background in ML systems, distributed computing, HPC or performance engineering, with experience working close to the infrastructure that runs modern ML models.
  • Strong Python and/or C++ skills are important, alongside a solid understanding of PyTorch and GPU computing.
  • Experience with technologies such as CUDA, NCCL or Triton would be highly relevant, as would hands-on work with inference frameworks including vLLM, TensorRT-LLM or SGLang.
  • You should be comfortable reasoning about performance at a systems level — understanding how GPU utilisation, memory bandwidth, communication overhead and model architecture interact to determine serving performance.

This is an opportunity to own meaningful parts of an inference stack being built from the ground up, rather than maintaining an established platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
AI Inference Engineer
AI Inference Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
Senior ML Inference Engineer: High-Performance GPU Systems
Senior ML Inference Engineer: High-Performance GPU Systems

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000