CloudDevs: Senior Site Reliability Engineer (SRE)

Breakout Tools

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves defining SLIs/SLOs, optimizing monitoring, and collaborating closely with engineering teams. Candidates should be comfortable writing production-grade code in Go, Python, or Node.js and thrive in fast-paced environments.

Qualifications

  • 5+ years in SRE, DevOps, or Platform Engineering roles.
  • Strong experience with cloud infrastructure (AWS preferred).
  • Deep knowledge of observability tools.
  • Strong debugging skills across services and networking.
  • Hands-on experience designing and monitoring SLIs/SLOs.

Responsibilities

  • Work as a hands-on engineer focused on system reliability.
  • Define and track SLIs, SLOs, and error budgets.
  • Optimize monitoring cost and signal quality.
  • Improve deployment safety and UAT pipelines.
  • Lead resilience work like failover drills and chaos tests.

Skills

Site Reliability Engineering
DevOps
Platform Engineering
Cloud Infrastructure
Observability Tools
Production-grade code in Go, Python, or Node.js

Tools

AWS
Terraform
Kubernetes
DataDog
Prometheus
GitHub Actions
Jenkins
ArgoCD

Job description

CloudDevs works with fast-moving, venture-backed startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will either be placed directly into one of our partner startups or added to our vetted SRE network for future projects.

This role is ideal for engineers who care about reliability, metrics, performance, and building simple, scalable systems. If you enjoy designing for scale and improving how teams ship software, you’ll fit right in.

Key Responsibilities
  • Work as a hands‑on engineer focused on system reliability, performance, and observability.
  • Define and track SLIs, SLOs, and error budgets.
  • Optimize monitoring cost and signal quality across metrics, logs, and traces.
  • Improve deployment safety, canary rollouts, and UAT pipelines.
  • Build tools for automated and local performance testing and track benchmarks.
  • Lead resilience work like failover drills, chaos tests, and redundancy checks.
  • Partner with engineering teams to improve scaling patterns and architecture as the product grows.
  • Support incident response processes and help reduce operational noise.
  • Write clean, maintainable code in Go, Python, or Node.js.
  • Contribute to CI/CD improvements and automation efforts.
  • Collaborate with engineers across teams to raise reliability standards.
Requirements
  • 5+ years in SRE, DevOps, or Platform Engineering roles.
  • Strong experience with cloud infrastructure (AWS preferred), Terraform, and Kubernetes.
  • Deep knowledge of observability tools like DataDog, Prometheus, or OpenTelemetry.
  • Strong debugging skills across services, networking, and data layers.
  • Hands‑on experience designing and monitoring SLIs/SLOs.
  • Experience with CI/CD tools such as GitHub Actions, Jenkins, or ArgoCD.
  • Ability to write production‑grade code in Go, Python, or Node.js.
  • Comfort working independently in fast‑paced environments.
Nice to Have
  • Experience tuning observability costs and optimizing data ingestion.
  • Exposure to chaos engineering and progressive deployments.
  • Background with high‑throughput or latency‑sensitive systems.
  • AWS at scale (EKS, Lambda, DynamoDB, S3).
  • Experience in regulated industries like fintech, payments, or SOC2 environments.
  • Performance testing pipelines or load‑testing automation.
  • Experience handling systems processing tens of millions of API calls.
Open Pool for SREs

Even if you don’t meet every requirement or aren’t a fit for the current role, strong SREs with real production experience are welcome to join our talent pool. We regularly place engineers with different strengths across reliability, DevOps, platform, observability, backend, and infrastructure engineering.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

Hybrid
USD 130,000 - 160,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000