Senior Real-Time Cloud Reliability Engineer

CloudZero

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CloudZero is seeking a Senior Site Reliability Engineer to own the reliability, performance, and observability of our real-time ingestion path. You’ll be a force multiplier for our engineering organization, enabling teams to ship features that help customers optimize cloud spend.

This is real infrastructure work at scale, with a serverless architecture and light on-call. You’ll shape reliability for multiple services and empower teams to operate confidently in production.

Qualifications

  • Strong production Python is your primary language and you own, test, and maintain it at scale.
  • You have defined an SLO yourself and know what not to alert on.
  • Experience operating asynchronous, event-driven systems and reasoning about back-pressure, lag, replay, poison messages, and partial failure (Kafka, Kinesis, SQS, Pulsar, or Step Functions).
  • You have driven reliability or platform changes across teams that didn’t report to you.
  • 5+ years building and operating distributed systems in AWS.
  • Infrastructure as Code in practice with CloudFormation/SAM or Terraform/Pulumi;
  • Hands-on experience instrumenting systems in monitoring tools like Datadog, Prometheus, Sumo Logic, or Splunk.

Responsibilities

  • Own the reliability practice for CloudZero's real-time ingestion path, including SLOs across team boundaries and architectural changes.
  • Sign off on shared critical paths before go-live and engage the pause conversation when an error budget is burned.
  • Instrument systems so failures surface quickly and debugging is data-driven.
  • Build observability into everything so you know about problems before customers do.
  • Design and maintain CloudFormation/SAM modules that provision reliable cloud resources; own infrastructure end-to-end.
  • Automate deployments, scaling, backups, and limit changes; replace repetitive work with systems.
  • Collaborate with Product Eng to design resilient services and drive adoption of best practices across 40+ engineers.
  • Drive cost and performance optimization, modeling CloudZero's own cloud usage as a case study.

Skills

Strong production Python
SLO definition
Event-driven systems
AWS distributed systems
Infrastructure as Code
Monitoring & observability
Debug under pressure
Explain technical issues
Breadth across areas
Chaos engineering / load testing

Tools

CloudFormation
SAM
Terraform
Pulumi
Datadog
Prometheus
Sumo Logic
GitHub Actions
Backstage/Cortex

Job description

CloudZero is seeking a Senior Site Reliability Engineer to own the reliability, performance, and observability of our real-time ingestion path. You’ll be a force multiplier for our engineering organization, enabling teams to ship features that help customers optimize cloud spend.

This is real infrastructure work at scale, with a serverless architecture and light on-call. You’ll shape reliability for multiple services and empower teams to operate confidently in production.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Real-Time Cloud Reliability Engineer
Senior Real-Time Cloud Reliability Engineer

cloudzero • Boston (MA)

On-site
USD 150,000 - 210,000
Senior CloudOps Engineer — Scale Reliability & Observability
Senior CloudOps Engineer — Scale Reliability & Observability

cloudzero • Boston (MA)

On-site
USD 100,000 - 130,000
Senior CloudOps Engineer
Senior CloudOps Engineer

CloudZero • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior CloudOps Engineer
Senior CloudOps Engineer

cloudzero • Boston (MA)

On-site
USD 150,000 - 210,000
Remote Senior Site Reliability Engineer - Cloud & Automation
Remote Senior Site Reliability Engineer - Cloud & Automation

Multi Media LLC • United States

On-site
USD 169,000 - 215,000
Fully Remote
Health Insurance
Vision Insurance
+10
Senior Cloud SRE & Reliability Engineer
Senior Cloud SRE & Reliability Engineer

Verygoodsecurity • United States

Hybrid
USD 110,000 - 140,000
Flexible work hours
Competitive health benefits
VGS stock options
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Platform Reliability Engineer (Remote/Hybrid)
Senior Platform Reliability Engineer (Remote/Hybrid)

Pantera Capital • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Senior Site Reliability Engineer: Scalable, Multi-Cloud
Senior Site Reliability Engineer: Scalable, Multi-Cloud

Movable Ink • United States

On-site
USD 184,000 - 240,000
Senior SRE: Database Infrastructure & Cloud Reliability
Senior SRE: Database Infrastructure & Cloud Reliability

PVH (Tommy Hilfiger/Calvin Klein) • Austin (TX)

On-site
USD 110,000 - 135,000