Site Reliability Engineer — Scale AI Infra with Ownership
Happyrobot Inc.
San Francisco (CA)
On-site
USD 100,000 - 140,000
Full time
14 days+
Application generator
Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Get past ATS filters
Benefits offered by this job
Competitive salary + equity
Ownership & autonomy in projects
Opportunity to work with top-tier engineers
Job summary
Happyrobot Inc. is looking for an Infrastructure Engineer in San Francisco, California. This role involves leading the stability and observability of systems while debugging complex issues as they arise. Candidates should have over 3 years of experience with production systems, strong skills in Go and Kubernetes, as well as familiarity with monitoring tools. Join us at a high-growth AI startup backed by top investors, where you will have ownership of projects and competitive compensation.
Qualifications
3+ years of experience debugging production systems.
Strong ability to dive into unfamiliar backend codebases.
Experience with observability and monitoring tools.
Responsibilities
Own stability, observability, and debugging workflows.
Untangle complex failures in real time.
Design tools that enhance operational resilience.
Skills
Debugging production systems
Problem-solving
Go programming
Kubernetes
Observability tools
Tools
Datadog
Prometheus
Sentry
Job description
Happyrobot Inc. is looking for an Infrastructure Engineer in San Francisco, California. This role involves leading the stability and observability of systems while debugging complex issues as they arise. Candidates should have over 3 years of experience with production systems, strong skills in Go and Kubernetes, as well as familiarity with monitoring tools. Join us at a high-growth AI startup backed by top investors, where you will have ownership of projects and competitive compensation.