Production ML Inference Engineer — Unlimited PTO

Thinking Machines Lab Inc.

San Francisco (CA)

On-site

USD 300,000 - 400,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health benefits
Dental & vision
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. is hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of serving our AI models to real users in a production environment.

You will bridge cutting-edge inference techniques with real traffic, rolling out new models safely and improving observability. The role focuses on multi-tenant serving, capacity planning, and incident response to keep the platform fast and resilient as usage grows.

Qualifications

  • Experience operating large-scale, latency-sensitive production systems.
  • Proficiency in Python and Go or another systems language.
  • Experience with observability, monitoring, and incident response for production services.
  • Strong understanding of distributed systems and how they fail at scale.

Responsibilities

  • Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform.
  • Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production.
  • Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly.
  • Partner with inference and research teams to productionize new serving techniques without compromising reliability.
  • Lead incident response for production inference issues, driving root cause analysis and durable fixes.
  • Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows.
  • Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale.

Skills

Python
Go or systems language
Observability
Distributed systems

Job description

Thinking Machines Lab Inc. is hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of serving our AI models to real users in a production environment.

You will bridge cutting-edge inference techniques with real traffic, rolling out new models safely and improving observability. The role focuses on multi-tenant serving, capacity planning, and incident response to keep the platform fast and resilient as usage grows.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Production ML Inference Engineer — Scale & Reliability
Production ML Inference Engineer — Scale & Reliability

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Inference Systems Engineer — Scalable, Reliable ML Serving
Inference Systems Engineer — Scalable, Reliable ML Serving

Speedrun Talent Network • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Software Engineer, Inference
Software Engineer, Inference

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Software Engineer, Inference
Software Engineer, Inference

Thinking Machines Lab Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health benefits
Dental & vision
Unlimited PTO
+2
Software Engineer, Inference
Software Engineer, Inference

Speedrun Talent Network • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
ML Infra Engineer - Scalable Training & Inference (Equity)
ML Infra Engineer - Scalable Training & Inference (Equity)

Snapchat • Palo Alto (CA)

On-site
USD 209,000 - 313,000
Senior ML Inference Engineer — Scale Production APIs
Senior ML Inference Engineer — Scale Production APIs

AssemblyAI, Inc. • New York (NY)

On-site
USD 190,000 - 225,000
Realtime ML Inference Engineer — Scalable Serving
Realtime ML Inference Engineer — Scalable Serving

Yobi AI • New York (NY)

Remote
USD 120,000 - 150,000
Competitive base salary
Meaningful equity
Annual bonus
+3
ML Platform Engineer: Inference & Production
ML Platform Engineer: Inference & Production

Xapply • San Francisco (CA)

On-site
USD 190,000 - 230,000
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1