Staff Site Reliability Engineer - Platform Resilience

Physics World

Santa Clara (CA)

Hybrid

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

IonQ, Inc. in Santa Clara, CA, is seeking a Staff Site Reliability Engineer to shape reliability across regions and services. You will own the reliability strategy, define standards, and mentor engineers while hands-on designing observability platforms and AI-driven triage workflows.

You will lead incident response, disaster recovery, and capacity planning, ensuring continuous service quality for cloud-managed SaaS with on-prem components. This role blends leadership with technical depth.

Qualifications

  • 7+ years of production engineering experience with reliability focus.
  • Hands-on experience operating large-scale fault-tolerant systems on AWS or GCP.
  • Observability ownership with SLOs, dashboards, and incident handling.
  • Incident command experience and ownership of root-cause analysis.
  • Proven ability to lead multiple teams and drive standards across services.
  • Experience with AI Ops, capacity planning, and automated remediation.

Responsibilities

  • Own production reliability end-to-end and set QoS standards.
  • Design and operate observability stack and governance of SLOs.
  • Define and manage incident response, including high-severity events.
  • Lead on-call rotations and escalation across regions and teams.
  • Drive disaster recovery planning and automated resilience improvements.
  • Collaborate with DevSecOps on cloud security posture and controls.
  • Mentor engineers and broaden impact across architecture and product teams.

Skills

Production engineering
Observability
Reliability engineering
Incident command
AWS/GCP
Multi-team leadership
AI Ops
Capacity planning
DevSecOps collaboration

Job description

IonQ, Inc. in Santa Clara, CA, is seeking a Staff Site Reliability Engineer to shape reliability across regions and services. You will own the reliability strategy, define standards, and mentor engineers while hands-on designing observability platforms and AI-driven triage workflows.

You will lead incident response, disaster recovery, and capacity planning, ensuring continuous service quality for cloud-managed SaaS with on-prem components. This role blends leadership with technical depth.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer: Platform Resilience & Observability
Staff Site Reliability Engineer: Platform Resilience & Observability

IonQ • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Lead Site Reliability Engineer — AI Ops & Resilience
Lead Site Reliability Engineer — AI Ops & Resilience

NightDragon Acquisition Corp. • Santa Clara (CA)

On-site
USD 152,000 - 228,000
Staff Reliability Engineer - Observability & AI Ops
Staff Reliability Engineer - Observability & AI Ops

IonQ • Santa Clara (CA)

Hybrid
USD 180,000 - 280,000
Staff SRE & AI-Driven Observability Architect
Staff SRE & AI-Driven Observability Architect

Physics World • Santa Clara (CA)

Hybrid
USD 170,000 - 250,000
Senior Service Reliability & AI Ops Engineer
Senior Service Reliability & AI Ops Engineer

NightDragon Acquisition Corp. • Santa Clara (CA)

On-site
USD 170,000 - 210,000
Home technology stipend
Medical, dental, vision coverage
401(k) matching
+1
Staff Site Reliability Engineer
Staff Site Reliability Engineer

IonQ • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Senior Staff Reliability & Observability Architect
Senior Staff Reliability & Observability Architect

Ionq • Santa Clara (CA)

On-site
USD 162,000 - 270,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Physics World • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Staff DevOps Engineer: Platform Automation & Security
Staff DevOps Engineer: Platform Automation & Security

Physics World • Santa Clara (CA)

Hybrid
USD 150,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+5
Staff Site Reliability Engineer — AI-Driven Reliability
Staff Site Reliability Engineer — AI-Driven Reliability

EarnIn • Mountain View (CA)

Hybrid
USD 252,000 - 308,000
Equity
Hybrid work model