A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves defining SLIs/SLOs, optimizing monitoring, and collaborating closely with engineering teams. Candidates should be comfortable writing production-grade code in Go, Python, or Node.js and thrive in fast-paced environments.
Qualifications
5+ years in SRE, DevOps, or Platform Engineering roles.
Strong experience with cloud infrastructure (AWS preferred).
Deep knowledge of observability tools.
Strong debugging skills across services and networking.
Hands-on experience designing and monitoring SLIs/SLOs.
Responsibilities
Work as a hands-on engineer focused on system reliability.
Define and track SLIs, SLOs, and error budgets.
Optimize monitoring cost and signal quality.
Improve deployment safety and UAT pipelines.
Lead resilience work like failover drills and chaos tests.
Skills
Site Reliability Engineering
DevOps
Platform Engineering
Cloud Infrastructure
Observability Tools
Production-grade code in Go, Python, or Node.js
Tools
AWS
Terraform
Kubernetes
DataDog
Prometheus
GitHub Actions
Jenkins
ArgoCD
Job description
A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves defining SLIs/SLOs, optimizing monitoring, and collaborating closely with engineering teams. Candidates should be comfortable writing production-grade code in Go, Python, or Node.js and thrive in fast-paced environments.