Head of ML Systems & Inference

Doist

San Francisco (CA)

On-site

USD 260,000 - 380,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, vision benefits
401(k) company match

Job summary

Inferact in San Francisco is seeking a Head of Engineering to build and lead the team developing systems that power vLLM and AI inference. You’ll ensure technical credibility at the inference layer, optimizing runtimes, memory and performance across GPUs and accelerators.

You’ll partner with founders to scale a senior-heavy organization, recruit rare ML systems talent, translate ambitious work into execution plans, and deliver high-performance inference across models, hardware, and deployment

Qualifications

  • Bachelor's degree or equivalent experience in CS/engineering or related field.
  • Engineering leadership experience building and scaling specialized teams in LLM inference or ML systems.
  • Deep technical credibility at the inference layer, including runtimes and hardware-software tradeoffs.

Responsibilities

  • Lead the engineering organization developing systems powering vLLM and Inferact.
  • Translate research into focused execution plans and measureable milestones.
  • Recruit, coach, and retain senior engineers and staff-level ICs for production readiness.

Skills

LLM Inference
Engineering Leadership
GPU optimization
Distributed Systems
People leadership

Education

Bachelor's degree

Tools

SGLang
PyTorch
Triton
XLA
ROCm

Job description

Inferact in San Francisco is seeking a Head of Engineering to build and lead the team developing systems that power vLLM and AI inference. You’ll ensure technical credibility at the inference layer, optimizing runtimes, memory and performance across GPUs and accelerators.

You’ll partner with founders to scale a senior-heavy organization, recruit rare ML systems talent, translate ambitious work into execution plans, and deliver high-performance inference across models, hardware, and deployment

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Engineering
Head of Engineering

Inferact • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
ML Inference & Systems Architect
ML Inference & Systems Architect

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Senior Backend Engineer, LLM Inference Systems
Senior Backend Engineer, LLM Inference Systems

Inception • San Francisco (CA)

On-site
USD 150,000 - 230,000
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides
ML Inference Systems Engineer
ML Inference Systems Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000