LLM Inference Systems Engineer

Acceler8 Talent

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

32 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Acceler8 Talent in San Francisco seeks an ML Inference Engineer to redesign the infrastructure behind its production AI platform. You will architect high-performance systems for serving LLMs, optimize inference for latency and throughput, and scale GPU workloads across a Kubernetes-based stack.

You will work close to the hardware and software stack, balancing trade-offs for impactful improvements at scale.

Qualifications

  • Strong background in ML systems, distributed computing, HPC or performance engineering.
  • Experience working close to infrastructure that runs modern ML models.
  • Proficiency in Python and/or C++.
  • Solid understanding of PyTorch and GPU computing.
  • Experience with CUDA, NCCL or Triton.
  • Hands-on work with inference frameworks such as vLLM, TensorRT-LLM or SGLang.

Responsibilities

  • Architect high-performance systems for serving LLMs at production scale.
  • Optimize inference for latency, throughput and cost efficiency.
  • Build distributed execution across multi-GPU and multi-node environments.
  • Improve GPU resource scheduling, allocation and utilization.
  • Profile the inference stack to identify bottlenecks across compute, memory and networking.
  • Scale GPU workloads across Kubernetes-based infrastructure.
  • Optimize techniques including batching, quantisation, KV caching and parallelism.

Skills

ML systems
Distributed computing
HPC
Performance engineering
Python
C++
PyTorch
GPU computing
CUDA
NCCL
Triton

Tools

CUDA Toolkit
NCCL
Triton

Job description

Acceler8 Talent in San Francisco seeks an ML Inference Engineer to redesign the infrastructure behind its production AI platform. You will architect high-performance systems for serving LLMs, optimize inference for latency and throughput, and scale GPU workloads across a Kubernetes-based stack.

You will work close to the hardware and software stack, balancing trade-offs for impactful improvements at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Systems Engineer
Senior ML Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Inference Engineer: High-Performance GPU Systems
Senior ML Inference Engineer: High-Performance GPU Systems

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
LLM Inference Platform Engineer
LLM Inference Platform Engineer

Baseten • New York (NY)

On-site
USD 216,000 - 360,000
Equity
Comprehensive health coverage
Flexible PTO
+4
Distributed LLM Inference Engineer - Scale & Resilience
Distributed LLM Inference Engineer - Scale & Resilience

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Equity
Health coverage
Flexible PTO
+4
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Senior LLM Inference Systems Engineer
Senior LLM Inference Systems Engineer

Snowflake • Menlo Park (CA)

On-site
USD 236,000 - 310,000
Medical insurance
Bonus & equity plan
401(k) retirement plan
+1
Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6