Site Reliability Engineer — Scale & Resilience for AI Ops

HappyRobot

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Top-Tier Compensation
Ownership & Autonomy
Opportunity to work at a high-growth startup

Job summary

A high-growth AI startup in San Francisco is seeking a Site Reliability Engineer to lead the scaling of operational resilience. In this role, you will own system stability and debugging workflows while tackling complex failures and enhancing proactive operations. Ideal candidates will have over 3 years of experience in debugging production systems, strong problem-solving skills, and familiarity with tools like Datadog and Prometheus. Join a dynamic team dedicated to redefining enterprise operations with cutting-edge AI technology.

Qualifications

  • 3+ years of hands-on experience debugging production systems.
  • Strong problem-solving skills and ability to dive into unfamiliar backend codebases.
  • Comfort with Python and Go for reading code and writing small tools.

Responsibilities

  • Own the stability, observability, and debugging workflows for systems.
  • Untangle complex failures in real time and design tools for clarity.
  • Shift from reactive to proactive operations.

Skills

Debugging production systems
Problem-solving skills
Python
Go
Communication under pressure

Tools

Datadog
Prometheus
Sentry

Job description

A high-growth AI startup in San Francisco is seeking a Site Reliability Engineer to lead the scaling of operational resilience. In this role, you will own system stability and debugging workflows while tackling complex failures and enhancing proactive operations. Ideal candidates will have over 3 years of experience in debugging production systems, strong problem-solving skills, and familiarity with tools like Datadog and Prometheus. Join a dynamic team dedicated to redefining enterprise operations with cutting-edge AI technology.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Scale & Observability
Site Reliability Engineer - Scale & Observability

gamma.app • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible work-from-home options
Creative and collaborative team environment
Site Reliability Engineer — Scale an AI‑Powered SaaS Platform
Site Reliability Engineer — Scale an AI‑Powered SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 140,000 - 165,000
Health insurance
Vision insurance
Dental plan
+2
Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Competitive compensation
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+3
Site Reliability Engineer — Scale AI Infra with Ownership
Site Reliability Engineer — Scale AI Infra with Ownership

Happyrobot Inc. • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary + equity
Ownership & autonomy in projects
Opportunity to work with top-tier engineers
Senior Site Reliability Engineer - Scale Resilient Systems
Senior Site Reliability Engineer - Scale Resilient Systems

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior SRE: Scale Reliability, Observability & CI/CD
Senior SRE: Scale Reliability, Observability & CI/CD

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,100 - 283,600
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
Staff Backend Engineer: AI-Driven SRE & Scalable Systems
Staff Backend Engineer: AI-Driven SRE & Scalable Systems

Resolve AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive Medical, Dental, and Vision Insurance
Monthly Housing Stipend
Flexible (Unlimited) Paid Time Off
+5
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000