AI-Native SRE: Cloud Infra for ML & Product Systems

Formation Bio

San Francisco (CA)

Hybrid

USD 186,000 - 232,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Formation Bio seeks a Senior Site Reliability Engineer to build and operate the infrastructure and delivery systems that power our AI-driven pharma platform. You will manage cloud workloads across AWS and other providers, ensure observability, and drive automation for product, data, and ML workloads.

You will partner with Product Engineering, Data Engineering, and Data Science to design scalable infrastructure, apply AI-assisted tooling, and enforce reliability across development to production

Qualifications

  • 5+ years in Site Reliability Engineering, infrastructure, or DevOps.
  • Production cloud experience with distributed systems and reliability focus.
  • Strong diagnostics, incident response, root cause analysis, observability, and automation.
  • Experience with AWS and Snowflake; Azure/GCP/Vercel is a plus.
  • Hands-on with Docker, GitHub, Kubernetes, Python, Terraform/OpenTofu, and networking.
  • Experience with ML/AI workloads and MLOps infrastructure a plus.
  • Familiar with COTS/FOSS software management and collaboration with teams.
  • Daily fluency with AI tools and LLMs; high validation standards.

Responsibilities

  • Own the shared infrastructure and operational platform for engineering workloads.
  • Build and operate secure, observable infrastructure for applications, data, ML pipelines.
  • Develop core AWS infrastructure and cloud outposts across dev/stage/prod.
  • Write and optimize infrastructure as code, CI/CD, and reusable platform patterns.
  • Establish SLOs, monitoring, runbooks, incident response, and post-incident reviews.
  • Collaborate with Product, Data, and AI teams to align architecture with needs.
  • Leverage AI tools to accelerate infra development and automation while validating outputs.
  • Participate in on-call rotations and balance automation with necessary manual steps.
  • Document requirements, design, and operating procedures; mentor engineers.

Skills

SRE experience
Observability
Incident response
Root cause analysis
Automation
AWS
Snowflake
Azure
GCP
Vercel
Docker
GitHub
Kubernetes
Python
Terraform/OpenTofu
Networking
Terragrunt
MLOps
Model serving
COTS/FOSS management
LLMs
AI tools
Collaboration
Regulated environment

Tools

Docker
GitHub
Kubernetes
Python
Terraform/OpenTofu
OpenTofu
Terragrunt

Job description

Formation Bio seeks a Senior Site Reliability Engineer to build and operate the infrastructure and delivery systems that power our AI-driven pharma platform. You will manage cloud workloads across AWS and other providers, ensure observability, and drive automation for product, data, and ML workloads.

You will partner with Product Engineering, Data Engineering, and Data Science to design scalable infrastructure, apply AI-assisted tooling, and enforce reliability across development to production

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE — AI-Native Cloud & Infra Leader
Senior SRE — AI-Native Cloud & Infra Leader

Formation Bio • San Francisco (CA)

Hybrid
USD 186,000 - 232,000
Equity
Comprehensive benefits
Generous perks
Senior AI-Native SRE Engineer
Senior AI-Native SRE Engineer

Formation • Town of Boston (NY)

Hybrid
USD 186,000 - 232,000
Equity
Benefits
AI-Native Infra & SRE Lead (Hybrid)
AI-Native Infra & SRE Lead (Hybrid)

Formation • Town of Boston (NY)

Hybrid
USD 186,000 - 232,000
AI-Driven Infra Engineering Manager — SRE Lead (Hybrid)
AI-Driven Infra Engineering Manager — SRE Lead (Hybrid)

Scorpion Therapeutics • San Francisco (CA)

Hybrid
USD 170,000 - 210,000
Hybrid work model
Engineering Manager, AI Infrastructure & SRE
Engineering Manager, AI Infrastructure & SRE

Formation Bio • New York (NY)

Hybrid
USD 186,000 - 232,000
Engineering Manager: AI-Driven Infrastructure & SRE
Engineering Manager: AI-Driven Infrastructure & SRE

Formation Bio • San Francisco (CA)

Hybrid
USD 186,000 - 232,000
Equity
Comprehensive benefits
Hybrid work model (3 days in office)
Senior SRE: AI-Driven Reliability & Cloud Automation
Senior SRE: AI-Driven Reliability & Cloud Automation

NDEAVOUR CONSULTING • United States

Hybrid
USD 120,000 - 150,000
Remote Office
Parking Space
Fun Office Space
+7
Engineering Manager, Infrastructure
Engineering Manager, Infrastructure

Scorpion Therapeutics • San Francisco (CA)

Hybrid
USD 170,000 - 210,000
Hybrid work model
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7
AI-Driven SRE Engineer for Cloud & Automation
AI-Driven SRE Engineer for Cloud & Automation

Skill • Austin (TX)

On-site
USD 140,000 - 190,000
Subsidized health plan
Retirement plan with match
Paid sick leave