Senior Staff Site Reliability Engineer

Pivotal Health

Santa Monica (CA)

Hybrid

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Health, dental, and vision coverage
401(k) retirement plan
Flexible time off
Company events

Job summary

Pivotal Health, a healthcare technology platform, is seeking a Staff Site Reliability Engineer to embed reliability and operational excellence into our production systems. You will lead hands-on engineering across cloud, observability, incident response, and automation, partnering with software, data, and security teams to scale a healthcare-grade platform with strong security and auditability.

You will collaborate with cross-functional teams to set strategy, implement resilient infrastructure,

Qualifications

  • 8+ years of experience in site reliability engineering, infrastructure engineering, platform engineering, or operating large-scale production systems.
  • Deep knowledge of cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, and modern deployment practices.
  • Experience designing, building, and operating highly available systems in a fast-growing production environment.
  • Strong understanding of observability, service-level objectives, capacity planning, incident management, disaster recovery, and performance engineering; Hands-on and technically credible, with the ability to debug complex issues across application, infrastructure, network, and data layers.
  • Experienced in creating automation and internal tooling that reduces operational toil and improves developer productivity.
  • Comfortable influencing architecture and engineering practices across teams without relying on formal authority.
  • A thoughtful communicator who can translate operational risk and technical tradeoffs for engineering, product, security, and business stakeholders; Pragmatic about balancing reliability, delivery speed, complexity, and cost.
  • Comfortable operating in ambiguity and building from scratch—you’re energized by greenfield work, not slowed down by it.

Responsibilities

  • Set Pivotal’s reliability strategy: Define the technical vision and roadmap for reliability, availability, scalability, and operational readiness. Establish clear service-level objectives and help teams make informed tradeoffs between reliability, velocity, and cost.
  • Design resilient, scalable infrastructure: Guide and implement improvements to our cloud architecture, deployment systems, networking, compute, storage, and other shared infrastructure. Ensure our systems can scale with increasing product usage, data volume, and workflow complexity.
  • Build world-class observability: Develop a cohesive approach to metrics, logs, traces, dashboards, and alerting. Give engineers the visibility they need to understand system behavior, identify emerging issues, and resolve production incidents quickly.
  • Improve incident response and resilience: Establish effective incident management, on-call, postmortem, and disaster recovery practices. Lead the response to complex incidents and ensure lessons result in durable improvements to our systems and processes.
  • Reduce operational toil through automation: Identify recurring manual work and build systems, tooling, and automation that make operating Pivotal’s platform safer and more efficient. Improve deployment workflows, capacity management, infrastructure provisioning, and production diagnostics.
  • Embed reliability across engineering: Partner with software, data, and AI engineers to improve system design, production readiness, and failure handling. Create reusable patterns, tooling, and standards that allow teams to build reliable services without becoming dependent on a centralized operations function.
  • Strengthen security and compliance: Work closely with security and compliance stakeholders to protect sensitive healthcare and financial data. Help ensure our infrastructure, access controls, audit trails, and operational practices meet applicable regulatory and customer requirements.
  • Provide technical leadership: Serve as a trusted technical partner to senior engineers and engineering leaders. Lead architecture reviews, mentor engineers, clarify complex tradeoffs, and raise the standard for infrastructure and operational engineering across the organization.

Skills

Site reliability engineering
Cloud infrastructure
Distributed systems
Observability
Automation
Incident management
Security & compliance awareness
Kubernetes
Terraform
AWS
GCP

Tools

Kubernetes
Terraform
AWS
GCP

Job description

Pivotal Health, a healthcare technology platform, is seeking a Staff Site Reliability Engineer to embed reliability and operational excellence into our production systems. You will lead hands-on engineering across cloud, observability, incident response, and automation, partnering with software, data, and security teams to scale a healthcare-grade platform with strong security and auditability.

You will collaborate with cross-functional teams to set strategy, implement resilient infrastructure,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer: Scalable Health Platform
Senior Site Reliability Engineer: Scalable Health Platform

Socket.dev • Williamsburg (VA)

Hybrid
USD 180,000 - 240,000
Competitive compensation including equ
Health coverage
401(k) retirement plan
+2
Staff SRE - Remote-Optional, Scale & Reliability Leader
Staff SRE - Remote-Optional, Scale & Reliability Leader

Pivotal Health • New York (NY)

Hybrid
USD 180,000 - 240,000
Competitive compensation with equity
Full health, dental, vision coverage
401(k) retirement savings plan
+2
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Pivotal Health • New York (NY)

Hybrid
USD 180,000 - 240,000
Competitive compensation with equity
Full health, dental, vision coverage
401(k) retirement savings plan
+2
Senior Site Reliability Engineer - Cloud & Resilience
Senior Site Reliability Engineer - Cloud & Resilience

pointclickcare • United States

On-site
USD 130,000 - 195,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Pivotal Health • Santa Monica (CA)

Hybrid
USD 180,000 - 250,000
Competitive compensation
Health, dental, and vision coverage
401(k) retirement plan
+2
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Pivotal Health • Los Angeles (CA)

Hybrid
USD 230,000 - 260,000
Competitive compensation
Full health, dental, vision
401(k) plan
+2
Senior DevOps Engineer for Healthcare AI Platform SRE
Senior DevOps Engineer for Healthcare AI Platform SRE

Transformcap • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Flexible hours
+1
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Socket.dev • Williamsburg (VA)

Hybrid
USD 180,000 - 240,000
Competitive compensation including equ
Health coverage
401(k) retirement plan
+2
Senior Site Reliability Engineer - Healthcare Infra Equity
Senior Site Reliability Engineer - Healthcare Infra Equity

Enzo Health • Lehi (UT)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
401k & Insurance
+2
Remote Senior Site Reliability Engineer — Reliability Lead
Remote Senior Site Reliability Engineer — Reliability Lead

Priority Technology Holdings, Inc. • Alpharetta (GA)

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
Medical, dental, and vision coverage
+1