Senior Site Reliability Engineer — Scale & Observability

Pivotal Health

New York (NY)

Hybrid

USD 230,000 - 260,000

Full time

14 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Health, dental, vision
401(k)
Flexible time off
Company-wide events

Job summary

Pivotal Health is seeking a Staff Site Reliability Engineer to embed reliability, scalability, and operational excellence into our healthcare platform. This senior IC role influences engineering across teams, focusing on resilient, scalable systems and robust production practices.

You will collaborate with software, data, AI, security, and product groups to improve observability, reduce toil, and ensure regulatory alignment as we scale our platform.

Qualifications

  • 8+ years of experience in site reliability engineering, infrastructure engineering, platform engineering, or operating large-scale production systems.
  • Deep knowledge of cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, and modern deployment practices.
  • Experience designing, building, and operating highly available systems in a fast-growing production environment.
  • Strong understanding of observability, service-level objectives, capacity planning, incident management, disaster recovery, and performance engineering; Hands-on and technically credible, with the ability to debug complex issues across application, infrastructure, network, and data layers.
  • Experienced in creating automation and internal tooling that reduces operational toil and improves developer productivity.
  • Comfortable influencing architecture and engineering practices across teams without relying on formal authority
  • A thoughtful communicator who can translate operational risk and technical tradeoffs for engineering, product, security, and business stakeholders; Pragmatic about balancing reliability, delivery speed, complexity, and cost

Responsibilities

  • Set Pivotal’s reliability strategy: Define the technical vision and roadmap for reliability, availability, scalability, and operational readiness. Establish clear service-level objectives and help teams make informed tradeoffs between reliability, velocity, and cost.
  • Design resilient, scalable infrastructure: Guide and implement improvements to our cloud architecture, deployment systems, networking, compute, storage, and other shared infrastructure. Ensure our systems can scale with increasing product usage, data volume, and workflow complexity.
  • Build world-class observability: Develop a cohesive approach to metrics, logs, traces, dashboards, and alerting. Give engineers the visibility they need to understand system behavior, identify emerging issues, and resolve production incidents quickly.
  • Improve incident response and resilience: Establish effective incident management, on-call, postmortem, and disaster recovery practices. Lead the response to complex incidents and ensure lessons result in durable improvements to our systems and processes.
  • Reduce operational toil through automation: Identify recurring manual work and build systems, tooling, and automation that make operating Pivotal’s platform safer and more efficient. Improve deployment workflows, capacity management, infrastructure provisioning, and production diagnostics.
  • Embed reliability across engineering: Partner with software, data, and AI engineers to improve system design, production readiness, and failure handling. Create reusable patterns, tooling, and standards that allow teams to build reliable services without becoming dependent on a centralized operations function.
  • Strengthen security and compliance: Work closely with security and compliance stakeholders to protect sensitive healthcare and financial data. Help ensure our infrastructure, access controls, audit trails, and operational practices meet applicable regulatory and customer requirements.
  • Provide technical leadership: Serve as a trusted technical partner to senior engineers and engineering leaders. Lead architecture reviews, mentor engineers, clarify complex tradeoffs, and raise the standard for infrastructure and operational engineering across the organization.

Skills

Experience
Cloud infrastructure
Observability
Incident management
Automation
Security
Team influence

Tools

Kubernetes

Job description

Pivotal Health is seeking a Staff Site Reliability Engineer to embed reliability, scalability, and operational excellence into our healthcare platform. This senior IC role influences engineering across teams, focusing on resilient, scalable systems and robust production practices.

You will collaborate with software, data, AI, security, and product groups to improve observability, reduce toil, and ensure regulatory alignment as we scale our platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer: Scale & Observability
Site Reliability Engineer: Scale & Observability

Empower Retirement, LLC • Overland Park (KS)

Hybrid
USD 87,400 - 123,400
Medical, dental, vision and life ins."
401(k) with company match
Tuition reimbursement
+1
Senior Platform Engineer — Scale Reliable Cloud Infrastructure
Senior Platform Engineer — Scale Reliable Cloud Infrastructure

Pivotal Health • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Equity
Health, dental, and vision
401(k)
+1
Senior Site Reliability Engineer to Scale Healthtech
Senior Site Reliability Engineer to Scale Healthtech

Transform9 • Birmingham (AL)

On-site
USD 100,000 - 150,000
Health Care Plan
Retirement Plan
Paid Time Off
+3
Senior DevOps Engineer for Healthcare AI Platform SRE
Senior DevOps Engineer for Healthcare AI Platform SRE

Transformcap • Palo Alto (CA)

Hybrid
USD 170,000 - 220,000
Equity
Medical insurance
Flexible hours
+1
Senior Site Reliability Engineer – Remote, Impact & Automation
Senior Site Reliability Engineer – Remote, Impact & Automation

Midwest Startups • United States

On-site
USD 175,000 - 185,000
Market-leading medical, dental, and視on
Stock options
Premium-Tier Origin Financial Wellness
+6
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer — Platform & Observability
Senior Site Reliability Engineer — Platform & Observability

Jobtailor • North Carolina

On-site
USD 180,000 - 240,000
Senior Platform SRE: Scale & Reliability
Senior Platform SRE: Scale & Reliability

United States Digital Space LLC • United States

Hybrid
USD 103,000 - 162,000
Health insurance
Vacation and RTT
Mental health and coaching
+7
Senior Site Reliability Engineer: Scalable Infra & Observability
Senior Site Reliability Engineer: Scalable Infra & Observability

Early Warning • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Matching
Paid Time Off
+1
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Pivotal Health • New York (NY)

Hybrid
USD 230,000 - 260,000
Equity
Health, dental, vision
401(k)
+2