Inference Systems Engineer — Scalable, Reliable ML Serving

Speedrun Talent Network

San Francisco (CA)

On-site

USD 300,000 - 400,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Thinking Machines is hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. This production-facing role sits at the center of the company, bridging cutting-edge inference techniques with the day-to-day reality of serving real traffic.

You’ll operate and scale production inference systems, own model rollout processes, and collaborate with research teams to productionize new serving techniques while maintaining high

Qualifications

  • Experience operating large-scale, latency-sensitive production systems.
  • Proficiency in Python and Go or another systems language.
  • Experience with observability, monitoring, and incident response for production services.
  • Strong understanding of distributed systems and how they fail at scale.

Responsibilities

  • Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform.
  • Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production.
  • Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly.
  • Partner with inference and research teams to productionize new serving techniques without compromising reliability.
  • Lead incident response for production inference issues, driving root cause analysis and durable fixes.
  • Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows.
  • Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale.

Skills

Python
Go
Observability
Distributed systems
Incident response

Job description

Thinking Machines is hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. This production-facing role sits at the center of the company, bridging cutting-edge inference techniques with the day-to-day reality of serving real traffic.

You’ll operate and scale production inference systems, own model rollout processes, and collaborate with research teams to productionize new serving techniques while maintaining high

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Production ML Inference Engineer — Scale & Reliability
Production ML Inference Engineer — Scale & Reliability

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Production ML Inference Engineer — Unlimited PTO
Production ML Inference Engineer — Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health benefits
Dental & vision
Unlimited PTO
+2
Realtime ML Inference Engineer — Scalable Serving
Realtime ML Inference Engineer — Scalable Serving

Yobi AI • New York (NY)

Remote
USD 120,000 - 150,000
Competitive base salary
Meaningful equity
Annual bonus
+3
Software Engineer, Inference
Software Engineer, Inference

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Software Engineer, Inference
Software Engineer, Inference

Thinking Machines Lab Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health benefits
Dental & vision
Unlimited PTO
+2
Software Engineer, Inference
Software Engineer, Inference

Speedrun Talent Network • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Real-Time ML Inference Engineer for Scalable Serving
Real-Time ML Inference Engineer for Scalable Serving

Yobi • New York (NY)

Hybrid
USD 100,000 - 150,000
Competitive Base Salary
Meaningful equity
Annual performance bonus
+3
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Senior ML Systems Engineer - Scalable Inference
Senior ML Systems Engineer - Scalable Inference

Atlassian • Seattle (WA)

Hybrid
USD 206,000 - 269,000
Health and wellbeing resources
Paid volunteer days
ML Systems Engineer: Scalable Inference & Distributed Infra
ML Systems Engineer: Scalable Inference & Distributed Infra

Atlassian • Seattle (WA)

Remote
USD 178,000 - 233,000
Health & wellbeing
Volunteer days
Equity opportunity