Site Reliability Engineer

HappyRobot

San Francisco (CA)

On-site

USD 100,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary + equity
Ownership of projects
World-class team

Job summary

HappyRobot, located in San Francisco, is searching for an Infrastructure Engineer to enhance our operational resilience. You will manage debugging workflows and ensure system stability while minimizing incident load.

This is a strategic role aimed at improving developer focus and uptime in a fast‑paced, high‑growth AI startup environment. Join a team where your contributions will significantly impact the technology landscape.

Qualifications

  • 3+ years of hands‑on experience debugging production systems.
  • Strong problem‑solving skills and ability to dive into unfamiliar backend codebases.
  • Familiarity with observability and monitoring tools.

Responsibilities

  • Own the stability, observability, and debugging workflows.
  • Design tools to improve operational resilience.
  • Help shift operations from reactive to proactive.

Skills

Debugging production systems
Problem-solving
Go
Kubernetes
Observability tools

Tools

Datadog
Prometheus
Sentry

Job description

About HappyRobot

HappyRobot is the infrastructure for enterprises to build and orchestrate AI workforces. Our AI workers don't just communicate - they make decisions, take action, and run operations autonomously across voice, email, and enterprise systems. Born in Y Combinator (S23) and backed by a16z and Base10 with over $60M raised, we power critical operations for global enterprises worldwide.

Our platform is battle-tested in the most demanding environments - where AI has real consequences. We started in logistics, built our own voice stack, models, and orchestration layer from the ground up, and are now bringing that infrastructure to every enterprise that runs the real economy. Learn more about our vision in our manifesto.

About the Role

We're looking for an Infrastructure Engineer to take the lead on scaling our operational resilience as we grow. You’ll own the stability, observability, and debugging workflows that keep our systems running smoothly. You'll be the go‑to person for untangling complex failures in real time, designing tools that turn chaos into clarity, and helping us shift from reactive to proactive operations.

This is a high‑impact, high‑trust role where you’ll shape how reliability is done — reducing incident load, building internal tooling, and directly improving developer focus and system uptime. If you love getting to the root of hard problems and making systems (and teams) stronger, this is your moment.

Must‑Have
  • 3+ years of hands‑on experience debugging production systems (logs, traces, incidents, etc.)
  • Strong problem‑solving skills and ability to dive into unfamiliar backend codebases
  • Strong Go and Kubernetes experience.
  • Familiarity with observability and monitoring tools (e.g., Datadog, Prometheus, Sentry)
  • Clear, calm communication under pressure — especially during live incidents
Nice‑to‑Have
  • Experience working with distributed systems or services at scale
  • Built or maintained internal tooling for on‑call teams or reliability workflows
  • Familiarity with deployment pipelines, CI/CD, or infra‑as‑code
  • Experience improving system observability (e.g., custom metrics, traces, log pipelines)
Why Join Us?
  • Opportunity to work at a high‑growth AI startup, backed by top investors.
  • Fast Growth — Backed by a16z and YC, on track for double‑digit ARR.
  • Top‑Tier Compensation — Competitive salary + equity in a high‑growth startup.
  • Ownership & Autonomy — Take full ownership of projects and ship fast.
  • Work With the Best — Join a world‑class team of engineers and builders.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Happyrobot Inc. • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary + equity
Ownership & autonomy in projects
Opportunity to work with top-tier engineers
Senior Software Engineer - Full-Stack
Senior Software Engineer - Full-Stack

Happyrobot Inc. • San Francisco (CA)

Hybrid
USD 220,000 - 250,000
Healthcare coverage
Dental coverage
Vision coverage
+1
Infrastructure Engineer
Infrastructure Engineer

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-Tier Compensation
Ownership & Autonomy
Opportunity to work at a high-growth startup
Software Engineer - Backend
Software Engineer - Backend

Dormont Manufacturing Co • San Francisco (CA)

On-site
USD 180,000 - 220,000
Competitive salary + equity
Healthcare coverage
Dental coverage
+1
AI Deployment Strategist: Customer Solutions & Growth
AI Deployment Strategist: Customer Solutions & Growth

Happyrobot Inc. • New York (NY)

On-site
Deployment Strategist · New York
Deployment Strategist · New York

Happyrobot Inc. • New York (NY)

On-site
USD 90,000 - 130,000
Healthcare coverage
Dental coverage
Vision coverage
+1
Forward Deployed Engineer
Forward Deployed Engineer

HappyRobot • New York (NY)

On-site
USD 120,000 - 200,000
Opportunity to work at a high-growth AI startup
Ownership & Autonomy
Competitive salary + equity
+2
Deployment Strategist
Deployment Strategist

Happyrobot Inc. • New York (NY)

On-site
USD 150,000 - 200,000
Comprehensive healthcare coverage
Competitive salary + equity
Ownership on projects
Deployment Strategist
Deployment Strategist

HappyRobot • New York (NY)

On-site
USD 90,000 - 120,000
Healthcare
Dental
Vision coverage
+2
Deployment Strategist
Deployment Strategist

Dormont Manufacturing Co • San Francisco (CA)

On-site
USD 150,000 - 200,000
Healthcare coverage
Competitive salary and equity
Ownership in project outcomes