Staff Site Reliability Engineer

IonQ

Santa Clara (CA)

On-site

USD 163,430 - 213,972

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IonQ is seeking a Staff Site Reliability Engineer in Santa Clara, CA, to lead reliability across regions and services. You will own the reliability strategy, define standards, and mentor teams while staying hands‑on with observability, SLOs, and disaster‑recovery automation.

Responsibilities include running high‑severity incident response, building AI Ops workflows for triage and remediation, and guiding on-call rotations with global coverage.

Qualifications

  • 7+ years of production engineering with hands‑on reliability work.
  • Experience operating large‑scale fault‑tolerant systems on AWS or GCP.
  • Observability ownership: instrumented systems and governed objectives and budgets.
  • Disaster recovery exercises with real failover validation.
  • Personally commanded SEV1/SEV2 incidents and drove systemic fixes.

Responsibilities

  • Own service-level objectives and error budgets end to end, and influence architecture and scaling decisions.
  • Design and operate the observability stack ensuring full instrumentation across services.
  • Define and manage service-level objectives, review with service owners, and drive corrective action for budget consumption.
  • Design and execute resilience and chaos experiments; validate safeguards against failure modes.
  • Lead incident response as incident commander for high‑severity incidents, including security events.
  • Run on‑call rotations and escalations with reliable handoffs and coverage.
  • Own disaster-recovery testing and failover validation, turning findings into improvements.
  • Co‑own cloud security posture management and runtime vulnerability detection with DevSecOps.

Skills

Production engineering
Reliability engineering
AWS/GCP experience
Incident command
Observability
Leadership
Cloud platforms

Tools

AWS
GCP

Job description

Company Overview

IonQ, Inc [NYSE: IONQ] is the world’s leading quantum platform and merchant supplier—delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ’s newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting‑edge systems that have helped customers and partners—including Amazon Web Services and AstraZeneca—to achieve 20× performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025 the company achieved 99.99% two‑qubit gate fidelity, setting a world record in quantum computing performance.

Location

Santa Clara, CA (Travel up to 25%)

Job ID: 1739

The Role

We are seeking a Staff Site Reliability Engineer. As a Staff SRE Engineer you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands‑on by designing and operating observability platforms, defining and governing SLO programs, leading high‑severity incident response, building resilience and disaster‑recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self‑healing.

Responsibilities
  • Production reliability: own service‑level objectives, error budgets, and production reliability outcomes end to end, and represent reliability in architecture and scaling decisions.
  • Engineer observability: design and operate the observability stack so production services are fully instrumented and define the standards platform and application teams follow.
  • Govern SLOs and error budgets: define and manage service‑level objectives, run regular reviews with service owners, and drive corrective action when services consume error budgets unsafely.
  • Drive resilience: design and execute chaos experiments and validate that failure modes are covered by tested safeguards.
  • Lead incident response: define the incident process and serve as incident commander for the highest‑severity incidents, including security incidents within the coverage window.
  • Run on‑call and escalation: establish and manage rotations and escalation paths that provide continuous coverage with clean follow‑the‑sun handoffs.
  • Disaster recovery: own disaster‑recovery testing and failover validation against defined recovery objectives and turn exercise findings into architectural and operational improvements.
  • Cloud security posture: co‑own cloud security posture management, runtime vulnerability detection, and configuration‑compliance monitoring with DevSecOps.
  • Data, streaming, and AI Ops: own reliability of stateful and streaming services, capacity planning and rightsizing, and autonomous agents for triage, predictive alerting, remediation, and self‑healing.
  • Scale the team and broaden impact: mentor engineers at different seniority levels, set standards adopted across teams, and align Architecture, DevSecOps, Cloud Operations, and Product Development behind a shared reliability roadmap.
Requirements
  • 7+ years of production engineering experience with recent hands‑on reliability work.
  • Hands‑on, recent experience operating large‑scale, fault‑tolerant production systems on AWS or GCP.
  • Observability ownership: have instrumented production systems and governed service‑level objectives and error budgets, not only installed dashboards.
  • Resilience practice: have designed and executed failure experiments or disaster‑recovery exercises with real failover validation.
  • Incident command: have personally commanded serious SEV1/SEV2 incidents and driven root cause through to a systemic fix.
  • Demonstrated ownership of reliability outcomes with measurable results (e.g., availability, mean time to recovery, and error‑budget adherence).
  • Evidence of multi‑team technical leadership through standards, review, coaching, and mechanisms adopted beyond one service or team.
Preferred Qualifications
  • Proven production experience with cloud security posture management, runtime vulnerability detection, and workload protection across cloud and distributed environments.
  • Strong experience prioritizing risk using identity, workload, and exposure‑path context to focus remediation on issues that materially increase attack likelihood and operational impact.
  • Experience with autonomous remediation and self‑healing workflows powered by AIOps, including Amazon Bedrock Agent Core or equivalent agentic automation frameworks.
  • Hands‑on experience in capacity management, resource rightsizing, efficiency engineering, and practical cost optimization based on FinOps principles.
  • Experience with load‑balancing design and operations, including health‑based failover, global traffic management, and performance optimization for highly available services.
  • Experience with AI traffic management via an LLM gateway, including request routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and observability for multi‑model or multi‑provider environments.
  • Ability to connect networking, security, and reliability considerations into cohesive platform design decisions that improve resilience, performance, and operability.
Compensation

The approximate base salary range for this position is $163,430 – $213,972. The total compensation package includes base, bonus, equity, and a range of benefit options found on our career site.

Benefits

Our benefits include comprehensive medical, dental, and vision plans; matching 401(k); unlimited PTO and paid holidays; parental/adoption leave; legal insurance; and a home technology stipend.

Equal Opportunity Employer

At IonQ, we believe in fair treatment, access, opportunity, and advancement for all while striving to identify and eliminate barriers. We empower employees to thrive by fostering a culture of autonomy, productivity, and respect. We are dedicated to creating an environment where individuals can feel welcomed, respected, supported, and valued. We are committed to equity and justice. We welcome different voices and viewpoints and do not discriminate on the basis of race, religion, ancestry, physical and/or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, transgender status, age, sexual orientation, military or veteran status, or any other basis protected by law. We are proud to be an Equal Employment Opportunity employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff Service Reliability and Operational Intelligence Engineer
Senior Staff Service Reliability and Operational Intelligence Engineer

Ionq • Santa Clara (CA)

On-site
USD 162,000 - 270,000
Senior Staff Service Reliability and Operational Intelligence Engineer New Santa Clara, California, United States
Senior Staff Service Reliability and Operational Intelligence Engineer New Santa Clara, California, United States

IonQ, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 188,000 - 270,000
Medical plan
Dental plan
Vision plan
+5
Staff DevOps Engineer
Staff DevOps Engineer

Ionq • Santa Clara (CA)

On-site
USD 187,944 - 246,068
Senior Research Software Engineer
Senior Research Software Engineer

IonQ • Boston (MA)

On-site
USD 145,000 - 192,000
Comprehensive medical, dental, and vision plans
Matching 401(k)
Unlimited PTO and paid holidays
+2
Senior Staff DevOps Engineer
Senior Staff DevOps Engineer

IonQ • Santa Clara (CA)

On-site
USD 216,000 - 283,000
Home‑technology stipend
401(k) matching
Unlimited PTO & holidays
+1
Senior Distributed Systems Engineer
Senior Distributed Systems Engineer

IonQ • Santa Clara (CA)

On-site
USD 163,000 - 214,000
Senior Staff Distributed Systems Engineer
Senior Staff Distributed Systems Engineer

IonQ • Santa Clara (CA)

On-site
USD 216,000 - 283,000
Staff Distributed Systems Engineer
Staff Distributed Systems Engineer

Ionq • Santa Clara (CA)

On-site
USD 187,944 - 246,068
Senior Physicist - Resilience Engineering & Operations
Senior Physicist - Resilience Engineering & Operations

IonQ • Bothell (WA)

Hybrid
USD 114,000 - 151,000
Quantum Success Engineer
Quantum Success Engineer

IonQ • United States

Hybrid
USD 128,000 - 183,000