Site Reliability Engineer

happyrobot.ai

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary
Equity

Job summary

HappyRobot is looking for a Site Reliability Engineer to scale operational resilience and own stability, observability, and debugging workflows across our AI infrastructure.

You will untangle complex failures in real time, design internal tools to turn chaos into clarity, and help shift from reactive to proactive operations, while collaborating with a world-class engineering team.

Qualifications

  • 3+ years of hands-on experience debugging production systems (logs, traces, incidents).
  • Strong problem-solving skills and ability to dive into unfamiliar backend codebases.
  • Strong Go and Kubernetes experience.
  • Familiarity with observability and monitoring tools (Grafana, Prometheus, Sentry).
  • Clear, calm communication under pressure - especially during live incidents.

Responsibilities

  • Own stability, observability, and debugging workflows.
  • Untangle complex failures in real time.
  • Design tools that turn chaos into clarity for developers.
  • Improve reliability and reduce incident load.
  • Shift operations from reactive to proactive.

Skills

Production debugging
Go
Kubernetes
Observability tools
Communication under pressure

Tools

Grafana
Prometheus
Sentry

Job description

About HappyRobot

HappyRobot is the infrastructure for enterprises to build and orchestrate AI workforces. Our AI workers don't just communicate through voice and email - they make decisions, take action, and run operations autonomously across entire enterprise systems. Born in Y Combinator (S23) and backed by a16z, Base10, Prysm Capital and Eurazeo with over $150M raised, we power critical operations for global enterprises worldwide.

Our platform is battle-tested in the most demanding environments, where AI has real consequences. We started in logistics, built our own voice stack, models, and orchestration layer from the ground up, and are now bringing that infrastructure to every enterprise that runs the real economy. Learn more about our vision in our manifesto.

About the Role

We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You’ll own the stability, observability, and debugging workflows that keep our systems running smoothly. You'll be the go-to person for untangling complex failures in real time, designing tools that turn chaos into clarity, and helping us shift from reactive to proactive operations.

This is a high-impact, high-trust role where you’ll shape how reliability is done - reducing incident load, building internal tooling, and directly improving developer focus and system uptime. If you love getting to the root of hard problems and making systems (and teams) stronger, this is your moment.

Must-Have
  • 3+ years of hands-on experience debugging production systems (logs, traces, incidents, etc.)

  • Strong problem-solving skills and ability to dive into unfamiliar backend codebases

  • Strong Go and Kubernetes experience.

  • Familiarity with observability and monitoring tools (e.g., Grafana, Prometheus, Sentry)

  • Clear, calm communication under pressure - especially during live incidents

Nice-to-Have
  • Experience working with distributed systems or services at scale

  • Built or maintained internal tooling for on-call teams or reliability workflows

  • Familiarity with deployment pipelines, CI/CD, or infra-as-code

  • Experience improving system observability (e.g., custom metrics, traces, log pipelines)

Why join us?
  • Join a world-class team of engineers and builders.

  • Backed by top investors including a16z, Y Combinator, Base10, Prysm Capital and Eurazeo.

  • Have ownership and autonomy of projects and are encouraged to ship.

  • Comprehensive Benefits including healthcare, dental, vision coverage.

  • Competitive salary + equity in a high-growth startup.

Our Operating Principles

Extreme Ownership

We take full responsibility for our work, outcomes, and team success. No excuses, no blame-shifting - if something needs fixing, we own it and make it better. This means stepping up, even when it's not "your job." If a ball is dropped, we pick it up. If a customer is unhappy, we fix it. If a process is broken, we redesign it. We don't wait for someone else to solve it - we lead with accountability and expect the same from those around us.

Craftsmanship

Putting care and intention into every task, striving for excellence, and taking deep ownership of the quality and outcome of your work. Craftsmanship means never settling for "just fine." We sweat the details because details compound. Whether it's a product feature, an internal doc, or a sales call - we treat it as a reflection of our standards. We aim to deliver jaw-dropping customer experiences by being curious, meticulous, and proud of what we build - even when nobody's watching.

We are 'majos' Be friendly & have fun with your coworkers. Always be genuine & honest, but kind. "Majo" is our way of saying: be a good human. Be approachable, helpful, and warm. We're building something ambitious, and it's easier (and more fun) when we enjoy the ride together. We give feedback with kindness, challenge each other with respect, and celebrate wins together without ego.

Urgency with Focus Create the highest impact in the shortest amount of time

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

HappyRobot • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary + equity
Ownership of projects
World-class team
Technical Customer Support Specialist
Technical Customer Support Specialist

HappyRobot • San Francisco (CA)

Hybrid
USD 110,000 - 170,000
Healthcare
Dental
Vision
+1
Forward Deployed Engineer
Forward Deployed Engineer

happyrobot.ai • New York (NY)

On-site
USD 120,000 - 170,000
Healthcare
Dental
Vision
+2
Software Engineer - Full-Stack
Software Engineer - Full-Stack

happyrobot.ai • San Francisco (CA)

On-site
USD 120,000 - 170,000
Healthcare
Dental
Vision
+1
Senior Software Engineer - Full-Stack
Senior Software Engineer - Full-Stack

happyrobot.ai • San Francisco (CA)

On-site
USD 150,000 - 210,000
Healthcare
Dental
Vision coverage
+1
PR & Communications Manager
PR & Communications Manager

happyrobot.ai • San Francisco (CA)

On-site
USD 130,000 - 190,000
Healthcare
Dental coverage
Vision coverage
+1
Product Operations Engineer
Product Operations Engineer

happyrobot.ai • San Francisco (CA)

On-site
USD 120,000 - 190,000
Salary + equity
Healthcare, dental, vision
Ownership & Autonomy
Site Reliability Engineer
Site Reliability Engineer

Happyrobot Inc. • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary + equity
Ownership & autonomy in projects
Opportunity to work with top-tier engineers
Technical Recruiting Manager
Technical Recruiting Manager

Happyrobot Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Healthcare benefits
Dental and vision coverage
Equity in a high-growth startup
PR & Communications Manager
PR & Communications Manager

Happyrobot Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Salary + equity
Healthcare, dental, vision
Ownership & autonomy