Staff SRE — Platform Reliability Lead

United States Digital Space LLC

United States

Hybrid

USD 127,000 - 161,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Deutschlandticket (Germany-wide public
28 vacation days
Work from abroad up to 10 days/year
Health insurance
Pension scheme (bAV) with 40% subsidy
DoctoGrowth program
Mental health & coaching services
Urban Sports Club membership
Hybrid/workplace policy
Subsidized meals
Relocation support
Access to AI tools and training

Job summary

United States Digital Space LLC is seeking a Staff Site Reliability Engineer to lead reliability and scalability across 170+ apps, supporting 520,000 health professionals and 90 million patients. You will mentor engineers, drive cross-functional reliability initiatives, and partner with product teams to embed resilience early in development.

The role blends infrastructure, developer experience, and product engineering, requiring strong backend skills (Go/Python/Ruby) and cloud expertise in

Qualifications

  • 8+ years in SRE, platform eng, or infra in large multi-team setups.
  • Cloud platforms (AWS, GCP, Azure) experience required.
  • Strong containerization/orchestration (Kubernetes) experience.
  • Implemented and operated SLIs, SLOs, and error budgets in production.
  • Experience managing on-call rotations and incident response.
  • Background in systems engineering with Go/Python/Ruby.
  • Proven leadership by influence across teams.
  • Fluent English in written and spoken form.

Responsibilities

  • Lead large-scale reliability initiatives across the platform including infra automation and observability.
  • Improve incident detection, response, and postmortem analysis capabilities.
  • Define and evolve SLOs, error budgets, and alerting standards across product teams.
  • Participate in on-call rotations and improve on-call experience and telemetry.
  • Mentor senior engineers and drive reliability engineering across the company.
  • Provide technical guidance to leadership and participate in architectural reviews.
  • Partner with software teams to embed reliability practices early in development lifecycle.

Skills

SRE experience 8+ years
Cloud platforms AWS/GCP/Azure
Kubernetes
SLIs/SLOs/Error budgets
On-call management
Backend programming (Go/Python/Ruby)
Leadership through influence
English fluency

Tools

Kubernetes
Terraform
AWS
GCP
Azure
Prometheus
OpenTelemetry
Datadog
ArgoCD

Job description

United States Digital Space LLC is seeking a Staff Site Reliability Engineer to lead reliability and scalability across 170+ apps, supporting 520,000 health professionals and 90 million patients. You will mentor engineers, drive cross-functional reliability initiatives, and partner with product teams to embed resilience early in development.

The role blends infrastructure, developer experience, and product engineering, requiring strong backend skills (Go/Python/Ruby) and cloud expertise in

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform SRE: Scale & Reliability
Senior Platform SRE: Scale & Reliability

United States Digital Space LLC • United States

Hybrid
USD 103,000 - 162,000
Health insurance
Vacation and RTT
Mental health and coaching
+7
Engineering Lead, Cloud Reliability & Platform (SRE)
Engineering Lead, Cloud Reliability & Platform (SRE)

United States Digital Space LLC • United States

Remote
USD 180,000 - 280,000
Annual learning budget
Health & wellness allowance
Home office budget
+5
Senior SRE Engineering Manager — Remote Reliability Leader
Senior SRE Engineering Manager — Remote Reliability Leader

United States Digital Space LLC • United States

Remote
USD 195,000 - 271,000
Competitive pay
Annual equity grants / ESPP
Health coverage
+8
Remote SRE Manager: Lead Reliability & Automation
Remote SRE Manager: Lead Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Site Reliability Engineer: Scale, Automate & Resilient Systems
Site Reliability Engineer: Scale, Automate & Resilient Systems

United States Digital Space LLC • United States

Remote
USD 120,000 - 210,000
Senior Federal SRE – Cloud Reliability & Automation
Senior Federal SRE – Cloud Reliability & Automation

United States Digital Space LLC • Washington

On-site
USD 170,000 - 260,000
Senior SRE - CI/CD Platform & AI-Driven Reliability
Senior SRE - CI/CD Platform & AI-Driven Reliability

United States Digital Space LLC • United States

Remote
USD 140,000 - 215,000
Market leader compensation
Comprehensive wellness programs
Paid vacation and holidays
+5
Senior Site Reliability Engineer - Automate & Scale
Senior Site Reliability Engineer - Automate & Scale

United States Digital Space LLC • New York (NY)

Hybrid
USD 179,000 - 226,000
Unlimited PTO
Employee stock options
Medical, dental, vision with HSA
+6
Remote SRE Lead — AI Reliability & Observability
Remote SRE Lead — AI Reliability & Observability

United States Digital Space LLC • United States

Remote
USD 198,000 - 303,000
Remote SRE Manager: Reliability & Scale for Fintech Health
Remote SRE Manager: Reliability & Scale for Fintech Health

NationsBenefits • Plantation (FL)

On-site
USD 140,000 - 190,000
Unlimited PTO
Fully remote work (US-based)