SRE for AI Training Pipelines & RL Runs

Mosaic.tech

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Relocation support
Unlimited PTO
Paid parental leave

Job summary

Thinking Machines is hiring a Site Reliability Engineer in San Francisco to own the health of post-training and RL pipelines, partnering with researchers to unblock training and harden infrastructure. Youll implement recovery, observability, and tooling to keep runs fast and reliable.

Youll operate in a high-scale ML environment, collaborate across teams, and participate in on-call rotations while driving permanent fixes and reducing toil for researchers.

Qualifications

  • 4+ years in production engineering or SRE for large-scale distributed systems.
  • Proven debugging across distributed systems including networks and schedulers.
  • Strong Python and/or Go/C++ programming skills with practical fixes.
  • Solid Linux internals and networking fundamentals.
  • Comfortable owning production systems and on-call rotations.

Responsibilities

  • Own reliability, performance, and uptime of large-scale post-training and RL jobs.
  • Collaborate with researchers during active model runs to unblock training.
  • Debug failures across accelerators, networking, storage, and schedulers.
  • Build monitoring, alerting, and automated recovery for self-healing runs.
  • Improve checkpointing, fault tolerance, and job scheduling efficiency.
  • Develop internal tools to reduce toil and improve cluster utilization.
  • Participate in on-call rotation and write postmortems for fixes.

Skills

Production engineering
SRE / site reliability
Python
Go/C++
Linux networking
On-call ownership
Debugging distributed systems

Tools

PyTorch
Ray
Kubernetes
Slurm
NCCL
InfiniBand

Job description

Thinking Machines is hiring a Site Reliability Engineer in San Francisco to own the health of post-training and RL pipelines, partnering with researchers to unblock training and harden infrastructure. Youll implement recovery, observability, and tooling to keep runs fast and reliable.

Youll operate in a high-scale ML environment, collaborate across teams, and participate in on-call rotations while driving permanent fixes and reducing toil for researchers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Training Infra SRE — Fast, Reliable RL Pipelines
ML Training Infra SRE — Fast, Reliable RL Pipelines

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Site Reliability Engineer, Post Training
Site Reliability Engineer, Post Training

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Visa sponsorship
Relocation support
Unlimited PTO
+1
SRE for AI Platform: Reliability at Scale
SRE for AI Platform: Reliability at Scale

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
SRE: Scalable ML Infra & CI/CD Architect
SRE: Scalable ML Infra & CI/CD Architect

Baseten • San Francisco (CA)

On-site
USD 165,000 - 330,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
SRE for Scalable AI Platform Reliability & Incident Response
SRE for Scalable AI Platform Reliability & Incident Response

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health/Dental/Vision
Unlimited PTO
Parental leave
+1
SRE: AI Inference Platform & ML Systems
SRE: AI Inference Platform & ML Systems

Cohere • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
SRE for AI Platform & ML Inference Infra
SRE for AI Platform & ML Inference Infra

Cohere • New York (NY)

Hybrid
USD 140,000 - 200,000
Weekly lunch stipend
Health and dental benefits
RRSP matching
+3
Site Reliability Engineer
Site Reliability Engineer

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
SRE, Applied ML: Scale, Monitor, Automate
SRE, Applied ML: Scale, Monitor, Automate

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+6