Site Reliability Engineer — Scale AI Infra with Ownership

Happyrobot Inc.

San Francisco (CA)

On-site

USD 100,000 - 140,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary + equity
Ownership & autonomy in projects
Opportunity to work with top-tier engineers

Job summary

Happyrobot Inc. is looking for an Infrastructure Engineer in San Francisco, California. This role involves leading the stability and observability of systems while debugging complex issues as they arise. Candidates should have over 3 years of experience with production systems, strong skills in Go and Kubernetes, as well as familiarity with monitoring tools. Join us at a high-growth AI startup backed by top investors, where you will have ownership of projects and competitive compensation.

Qualifications

  • 3+ years of experience debugging production systems.
  • Strong ability to dive into unfamiliar backend codebases.
  • Experience with observability and monitoring tools.

Responsibilities

  • Own stability, observability, and debugging workflows.
  • Untangle complex failures in real time.
  • Design tools that enhance operational resilience.

Skills

Debugging production systems
Problem-solving
Go programming
Kubernetes
Observability tools

Tools

Datadog
Prometheus
Sentry

Job description

Happyrobot Inc. is looking for an Infrastructure Engineer in San Francisco, California. This role involves leading the stability and observability of systems while debugging complex issues as they arise. Candidates should have over 3 years of experience with production systems, strong skills in Go and Kubernetes, as well as familiarity with monitoring tools. Join us at a high-growth AI startup backed by top investors, where you will have ownership of projects and competitive compensation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — Scale & Resilience for AI Ops
Site Reliability Engineer — Scale & Resilience for AI Ops

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Site Reliability Engineer
Site Reliability Engineer

HappyRobot • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary + equity
Ownership of projects
World-class team
Site Reliability Engineer — Build Scalable AI Infra
Site Reliability Engineer — Build Scalable AI Infra

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 140,000
Flexible work environment
Competitive salary
Growth trajectory
Site Reliability Engineer, Cloud Infra for AI Platform
Site Reliability Engineer, Cloud Infra for AI Platform

Anyscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Site Reliability Engineer, AI Infra & Observability
Site Reliability Engineer, AI Infra & Observability

Sierra • California (MO)

On-site
USD 150,000 - 210,000
Unlimited PTO
Medical, dental, vision
Retirement plan
+4
Senior AI Infra Engineer, On-Site in SF
Senior AI Infra Engineer, On-Site in SF

The Recruiting Guy • San Francisco (CA)

On-site
USD 175,000 - 250,000
Senior AI Infra Engineer — Scale Systems, Equity & PTO
Senior AI Infra Engineer — Scale Systems, Equity & PTO

King River Capital Group • San Francisco (CA)

On-site
USD 100,000 - 130,000
Great health insurance, dental, and vision
Gym and workspace stipends
Unlimited PTO
Site Reliability Engineer — ML Infra & Observability
Site Reliability Engineer — ML Infra & Observability

Baseten • San Francisco (CA)

On-site
USD 135,000 - 285,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Platform Engineer: Scale AI Runtime & Infra
Platform Engineer: Scale AI Runtime & Infra

Anything • San Francisco (CA)

On-site
USD 120,000 - 160,000