Senior Real-Time Cloud Reliability Engineer

cloudzero

Boston (MA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CloudZero is seeking a Senior Site Reliability Engineer to own the reliability, performance, and observability of its real-time ingestion path. You will drive SLOs, instrument systems with data-driven debugging, and automate deployments across a scalable, serverless architecture.

This role focuses on building reliable infrastructure, reducing incident impact, and enabling product teams to ship features that optimize customer cloud spend. On-call rotates lightly as CloudZero scales.

Qualifications

  • Strong production Python as primary language, scale-tested.
  • Own reliability outcomes with ownership beyond tasks.
  • 5+ years building and operating distributed systems in AWS.
  • Experience with asynchronous, event-driven systems and back-pressure.
  • Familiarity with observability and monitoring tools at scale.

Responsibilities

  • Own the reliability, performance, and observability of real-time ingestion path.
  • Sign off on critical paths, manage error budgets, and call pauses when needed.
  • Instrument systems to surface failures with data-driven debugging.
  • Build and integrate load generators, fault-injection, and SLO libraries.
  • Design CloudFormation and SAM modules for reliable cloud resources.
  • Automate deployments, scaling, backups, and limit changes across services.

Skills

Python
SRE
Distributed systems
Kafka/Kinesis
Observability
AWS
On-call ownership

Education

Bachelor's degree in Computer Science or equivalent

Tools

CloudFormation
SAM
Terraform
Pulumi
Datadog
Prometheus
Sumo Logic
Splunk
CloudWatch

Job description

CloudZero is seeking a Senior Site Reliability Engineer to own the reliability, performance, and observability of its real-time ingestion path. You will drive SLOs, instrument systems with data-driven debugging, and automate deployments across a scalable, serverless architecture.

This role focuses on building reliable infrastructure, reducing incident impact, and enabling product teams to ship features that optimize customer cloud spend. On-call rotates lightly as CloudZero scales.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Real-Time Cloud Reliability Engineer
Senior Real-Time Cloud Reliability Engineer

CloudZero • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior CloudOps Engineer — Scale Reliability & Observability
Senior CloudOps Engineer — Scale Reliability & Observability

cloudzero • Boston (MA)

On-site
USD 100,000 - 130,000
Senior Platform Reliability Engineer (Remote/Hybrid)
Senior Platform Reliability Engineer (Remote/Hybrid)

Pantera Capital • San Jose (CA)

Hybrid
USD 99,000 - 229,000
Senior CloudOps Engineer
Senior CloudOps Engineer

CloudZero • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior CloudOps Engineer
Senior CloudOps Engineer

cloudzero • Boston (MA)

On-site
USD 150,000 - 210,000
Senior SRE: Database Infrastructure & Cloud Reliability
Senior SRE: Database Infrastructure & Cloud Reliability

PVH (Tommy Hilfiger/Calvin Klein) • Austin (TX)

On-site
USD 110,000 - 135,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior/Staff CloudOps Engineer
Senior/Staff CloudOps Engineer

cloudzero • Boston (MA)

On-site
USD 100,000 - 130,000
Senior Site Reliability Engineer – Remote, Impact & Automation
Senior Site Reliability Engineer – Remote, Impact & Automation

Midwest Startups • United States

On-site
USD 175,000 - 185,000
Market-leading medical, dental, and視on
Stock options
Premium-Tier Origin Financial Wellness
+6
Remote Senior Site Reliability Engineer - Cloud & Automation
Remote Senior Site Reliability Engineer - Cloud & Automation

Multi Media LLC • United States

On-site
USD 169,000 - 215,000
Fully Remote
Health Insurance
Vision Insurance
+10