Production ML Inference Engineer — Scale & Reliability

AI Chopping Block, Inc.

San Francisco (CA)

On-site

USD 300,000 - 400,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Thinking Machines is hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. This production-facing role sits at the center of a fast-growing platform, bridging cutting-edge inference techniques with real traffic, including rollouts, capacity planning, incidents, and resiliency.

You will operate and scale live traffic systems, own model rollout processes, and collaborate with research teams to productionize new

Qualifications

  • Experience operating large-scale, latency-sensitive production systems.
  • Proficiency in Python and Go or another systems language.
  • Experience with observability, monitoring, and incident response for production services.
  • Strong understanding of distributed systems and how they fail at scale.

Responsibilities

  • Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform.
  • Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production.
  • Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly.
  • Partner with inference and research teams to productionize new serving techniques without compromising reliability.
  • Lead incident response for production inference issues, driving root cause analysis and durable fixes.
  • Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows.
  • Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale.

Skills

Large-scale systems
Python
Go
Observability
Distributed systems
Incident response

Job description

Thinking Machines is hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. This production-facing role sits at the center of a fast-growing platform, bridging cutting-edge inference techniques with real traffic, including rollouts, capacity planning, incidents, and resiliency.

You will operate and scale live traffic systems, own model rollout processes, and collaborate with research teams to productionize new

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Systems Engineer — Scalable, Reliable ML Serving
Inference Systems Engineer — Scalable, Reliable ML Serving

Speedrun Talent Network • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Production ML Inference Engineer — Unlimited PTO
Production ML Inference Engineer — Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health benefits
Dental & vision
Unlimited PTO
+2
Senior ML Inference Engineer — Scale Production APIs
Senior ML Inference Engineer — Scale Production APIs

AssemblyAI, Inc. • New York (NY)

On-site
USD 190,000 - 225,000
Production ML Systems Engineer - Scale & Observability
Production ML Systems Engineer - Scale & Observability

Careervitablr • United States

Remote
USD 140,000 - 210,000
Software Engineer, Inference
Software Engineer, Inference

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Software Engineer, Inference
Software Engineer, Inference

Thinking Machines Lab Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health benefits
Dental & vision
Unlimited PTO
+2
Software Engineer, Inference
Software Engineer, Inference

Speedrun Talent Network • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
ML Infra Engineer - Scalable Training & Inference (Equity)
ML Infra Engineer - Scalable Training & Inference (Equity)

Snapchat • Palo Alto (CA)

On-site
USD 209,000 - 313,000
Production ML Engineer — Real‑Time Safety & Scale
Production ML Engineer — Real‑Time Safety & Scale

Cinder Technologies • New York (NY)

On-site
USD 150,000 - 230,000
Health benefits
Vision & dental
401(k) match
+3
Senior ML Inference Engineer — Production Systems
Senior ML Inference Engineer — Production Systems

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000