SRE for Scalable AI Platform Reliability & Incident Response

Thinking Machines Lab

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

13 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health/Dental/Vision
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to drive the reliability of Tinker end-to-end, collaborating with platform engineers and researchers to make every layer of the system robust.

The role emphasizes defining end-to-end reliability, building observability, incident response, and multi-tenant isolation for large-scale distributed training workloads on Kubernetes.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Experience in distributed systems, cloud infrastructure, or site reliability engineering.
  • Proficiency writing software to solve reliability problems, including building tooling and automation.
  • Experience with production incident response, postmortems, and systematic reliability improvement.
  • Strong communication skills and track record of coordination across engineering and research teams.

Responsibilities

  • Define and own end-to-end reliability, from CI/CD flows to production observability and incident response.
  • Develop appropriate Service Level Objectives for distributed training systems, balancing job completion reliability and scheduling latency with development velocity.
  • Design and implement monitoring and observability across the full training path.
  • Drive incident response for Tinker platform issues, ensuring rapid recovery, thorough incident reviews, and systematic improvements that prevent recurrence.
  • Harden multi-tenant isolation and resource scheduling so that LoRA-based workload co-scheduling maximizes utilization without compromising reliability or data separation.
  • Collaborate with security teams to address production vulnerabilities.

Skills

Distributed systems
Cloud infrastructure
Software tooling
Incident response
Cross-team communication

Education

Bachelor's degree or equivalent in CS/Engineering

Tools

Kubernetes
CI/CD tooling
Observability tooling

Job description

Thinking Machines Lab is seeking a Site Reliability Engineer (SRE) to drive the reliability of Tinker end-to-end, collaborating with platform engineers and researchers to make every layer of the system robust.

The role emphasizes defining end-to-end reliability, building observability, incident response, and multi-tenant isolation for large-scale distributed training workloads on Kubernetes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE for AI Platform: Reliability at Scale
SRE for AI Platform: Reliability at Scale

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Site Reliability Engineer, Production
Site Reliability Engineer, Production

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health/Dental/Vision
Unlimited PTO
Parental leave
+1
Senior SRE, MLOps & AI Infra Architect
Senior SRE, MLOps & AI Infra Architect

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,000 - 284,000
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
Site Reliability Engineer, Production
Site Reliability Engineer, Production

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Founding SRE — Build Reliable, Scalable Systems
Founding SRE — Build Reliable, Scalable Systems

Incident IQ • Atlanta (GA)

On-site
USD 120,000 - 190,000
Medical benefits
Dental benefits
Vision benefits
+3
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
SRE for AI Training Pipelines & RL Runs
SRE for AI Training Pipelines & RL Runs

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Visa sponsorship
Relocation support
Unlimited PTO
+1
Senior Data Infra SRE: Reliability, Scale & Automation
Senior Data Infra SRE: Reliability, Scale & Automation

TikTok • Seattle (WA)

On-site
USD 207,000 - 368,000
Senior SRE — AI Platform Reliability & Automation
Senior SRE — AI Platform Reliability & Automation

nscaleoperationsukltd • Houston (TX)

On-site
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work style