Staff Site Reliability Engineer

IonQ

Santa Clara (CA)

Hybrid

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

IonQ, Inc. is seeking a Staff Site Reliability Engineer to lead reliability for cloud-managed SaaS products, with a focus on observability, SLOs, and disaster recovery.

Based in Santa Clara, CA, with remote-work potential a few days per week, you will own the reliability strategy and drive improvements across regions and services. You will design and operate observability platforms, mentor teams, and strengthen incident response while co-owning cloud security posture and resilience initiatives

Qualifications

  • Experience in reliability engineering for cloud-based SaaS platforms.

Responsibilities

  • Production reliability — own service-level objectives, error budgets, and production reliability outcomes end to end.
  • Engineer observability — design and operate the observability stack so production services are fully instrumented and define the standards platform and application.

Skills

Production reliability
Observability
Incident response
Chaos engineering
Disaster recovery
Cloud security

Job description

About IonQ:

IonQ, Inc. [NYSE: IONQ] is the world's leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ's newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting-edge systems that have been helping customers and partners including Amazon Web Services, and AstraZeneca achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two-qubit gate fidelity, setting a world record in quantum computing performance.

About IonQ:

IonQ, Inc. [NYSE: IONQ] is the world's leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ's newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting-edge systems that have been helping customers and partners including Amazon Web Services, and AstraZeneca achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two-qubit gate fidelity, setting a world record in quantum computing performance. Headquartered in College Park, Maryland, IonQ has operations in California, Colorado, Massachusetts, Tennessee, Washington, Italy, South Korea, Sweden, Switzerland, Canada, and the United Kingdom. Our quantum computing services are available through all major cloud providers, while we also meet the needs of networking and sensing customers across land, sea, air, and space. IonQ is making quantum platforms more accessible and impactful than ever before.

Location:

This role is based at our Santa Clara, CA office, with the option to work a few days a week remotely.

Travel:

Up to 25%

Job ID:

1874

The Role:

The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites.

The Site Reliability Engineering discipline keeps the platform stable and reliable, with a strong focus on service continuity and customer experience. It owns production reliability, service-level objectives, observability architecture, backup and disaster recovery, incident response, and resilience, and co-owns cloud security posture and runtime vulnerability management with DevSecOps.

As Staff Site Reliability Engineer, you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing.

The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system.

  • Production reliability, SLOs, and error budgets — the reliability of production services end to end, including standards, governance, and escalation for Tier-1 and Tier-2 services.
  • Observability architecture and standards — metrics, logs, distributed tracing, and profiles instrumented across production systems, with consistent platform-wide standards.
  • Chaos engineering and resilience — failure-injection experiments and validation of recovery mechanisms in pre-production and production environments.
  • Backup and disaster recovery — backup validation, disaster-recovery architecture, failover testing, and recovery verification against defined RTO and RPO objectives.
  • Cloud security posture — cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring, co-owned with DevSecOps.
  • Data and streaming platform reliability — reliability engineering for Postgres, Redis/Valkey, Kafka, OpenSearch, and other critical stateful services.
  • Capacity, efficiency, and AI Ops — resource rightsizing, predictive alerting, autonomous triage, remediation automation, and self-healing workflows.
  • Incident response and command — severity classification, incident command, executive communication, and blameless post-incident review for the highest-severity events.
  • On-call and escalation — rotation design, operational readiness, escalation policy, and clean follow-the-sun handoffs across regions.
Responsibilities:
  • Production reliability — own service-level objectives, error budgets, and production reliability outcomes end to end, and represent reliability in architecture and scaling decisions.
  • Engineer observability — design and operate the observability stack so production services are fully instrumented and define the standards platform and application
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Physics World • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Staff Service Reliability and Operational Intelligence Engineer
Staff Service Reliability and Operational Intelligence Engineer

IonQ • Santa Clara (CA)

Hybrid
USD 180,000 - 280,000
Senior Staff Service Reliability and Operational Intelligence Engineer
Senior Staff Service Reliability and Operational Intelligence Engineer

Ionq • Santa Clara (CA)

On-site
USD 162,000 - 270,000
Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

Clutch Canada • Santa Clara (CA)

On-site
USD 150,000 - 190,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

NightDragon Acquisition Corp. • Santa Clara (CA)

Hybrid
USD 152,000 - 228,000
Staff Site Reliability Engineer - Platform Resilience
Staff Site Reliability Engineer - Platform Resilience

Physics World • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Staff Service Reliability and Operational Intelligence Engineer
Staff Service Reliability and Operational Intelligence Engineer

NightDragon Acquisition Corp. • Santa Clara (CA)

Hybrid
USD 170,000 - 210,000
Home technology stipend
Medical, dental, vision coverage
401(k) matching
+1
Staff Service Reliability and Operational Intelligence Engineer
Staff Service Reliability and Operational Intelligence Engineer

Physics World • Santa Clara (CA)

Hybrid
USD 170,000 - 250,000
Staff Site Reliability Engineer: Platform Resilience & Observability
Staff Site Reliability Engineer: Platform Resilience & Observability

IonQ • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Senior Distributed Systems Engineer
Senior Distributed Systems Engineer

Clutch Canada • Santa Clara (CA)

On-site
USD 180,000 - 240,000