Senior ML Inference Engineer — Production Systems

MakerMaker.AI

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability.

The ideal candidate will have 3+ years of experience in production-grade serving infrastructure, be fluent in Python, and have strong knowledge in GPU-accelerated inference. Excellent communication skills are essential as you'll be creating runbooks for incident resolution.

Qualifications

  • 3+ years building production-grade, large-scale serving infrastructure.
  • Experience shipping production infrastructure handling millions of requests.
  • Ability to read flame graphs and perform analytical changes.

Responsibilities

  • Build and operate production inference systems serving large models.
  • Own performance characteristics: throughput, latency, cost-per-token.
  • Diagnose production incidents and write systemic fixes.

Skills

Performance profiling and optimization fluency
Fluent Python
Strong distributed systems experience
Experience with GPU-accelerated inference at scale
Good written communication

Tools

C++
CUDA
ROCm
Triton

Job description

MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability.

The ideal candidate will have 3+ years of experience in production-grade serving infrastructure, be fluent in Python, and have strong knowledge in GPU-accelerated inference. Excellent communication skills are essential as you'll be creating runbooks for incident resolution.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer: Scale AI in Production
Senior ML Inference Engineer: Scale AI in Production

Oscar • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Equity
401k matching
Medical coverage
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
ML Inference Systems Engineer
ML Inference Systems Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Systems Engineer: Scale Training & Inference
ML Systems Engineer: Scale Training & Inference

Doist • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive cash compensation
Startup equity
ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer — On-Site in Palo Alto, High-Impact
ML Systems Engineer — On-Site in Palo Alto, High-Impact

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers
Distributed Systems Engineer - Data & Inference Platform
Distributed Systems Engineer - Data & Inference Platform

OpenTalent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Flexible work
Adaption Passport
Lunch Stipend
+1
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1