Staff Site Reliability Engineer

Physics World

Santa Clara (CA)

Hybrid

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

IonQ, Inc. in Santa Clara, CA, is seeking a Staff Site Reliability Engineer to shape reliability across regions and services. You will own the reliability strategy, define standards, and mentor engineers while hands-on designing observability platforms and AI-driven triage workflows.

You will lead incident response, disaster recovery, and capacity planning, ensuring continuous service quality for cloud-managed SaaS with on-prem components. This role blends leadership with technical depth.

Qualifications

  • 7+ years of production engineering experience with reliability focus.
  • Hands-on experience operating large-scale fault-tolerant systems on AWS or GCP.
  • Observability ownership with SLOs, dashboards, and incident handling.
  • Incident command experience and ownership of root-cause analysis.
  • Proven ability to lead multiple teams and drive standards across services.
  • Experience with AI Ops, capacity planning, and automated remediation.

Responsibilities

  • Own production reliability end-to-end and set QoS standards.
  • Design and operate observability stack and governance of SLOs.
  • Define and manage incident response, including high-severity events.
  • Lead on-call rotations and escalation across regions and teams.
  • Drive disaster recovery planning and automated resilience improvements.
  • Collaborate with DevSecOps on cloud security posture and controls.
  • Mentor engineers and broaden impact across architecture and product teams.

Skills

Production engineering
Observability
Reliability engineering
Incident command
AWS/GCP
Multi-team leadership
AI Ops
Capacity planning
DevSecOps collaboration

Job description

About IonQ

IonQ, Inc . [NYSE: IONQ] is the world's leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ's newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting‑edge systems that have been helping customers and partners including Amazon Web Services, and AstraZeneca achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two‑qubit gate fidelity, setting a world record in quantum computing performance.

Job Details

IonQ, Inc . [NYSE: IONQ] is the world's leading quantum platform and merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ's newest generation of quantum computers, the IonQ Tempo, is the latest in a line of cutting‑edge systems that have been helping customers and partners including Amazon Web Services, and AstraZeneca achieve 20x performance results and accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense. In 2025, the company achieved 99.99% two‑qubit gate fidelity, setting a world record in quantum computing performance.

Headquartered in College Park, Maryland, IonQ has operations in California, Colorado, Massachusetts, Tennessee, Washington, Italy, South Korea, Sweden, Switzerland, Canada, and the United Kingdom. Our quantum computing services are available through all major cloud providers, while we also meet the needs of networking and sensing customers across land, sea, air, and space. IonQ is making quantum platforms more accessible and impactful than ever before.

Location: This role is based at our Santa Clara, CA office, with the option to work a few days a week remotely.

Travel: Up to 25%

Job ID: 1874

The Role

The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud‑managed SaaS products with on‑premises components deployed at customer sites. The Site Reliability Engineering discipline keeps the platform stable and reliable, with a strong focus on service continuity and customer experience. It owns production reliability, service‑level objectives, observability architecture, backup and disaster recovery, incident response, and resilience, and co‑owns cloud security posture and runtime vulnerability management with DevSecOps. As Staff Site Reliability Engineer, you set the technical direction for reliability across regions and services. You own the reliability strategy, define the standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands‑on by designing and operating observability platforms, defining and governing SLO programs, leading high‑severity incident response, building resilience and disaster‑recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self‑healing. The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system.

  • Production reliability, SLOs, and error budgets - the reliability of production services end to end, including standards, governance, and escalation for Tier-1 and Tier-2 services.
  • Observability architecture and standards - metrics, logs, distributed tracing, and profiles instrumented across production systems, with consistent platform-wide standards.
  • Chaos engineering and resilience - failure-injection experiments and validation of recovery mechanisms in pre‑production and production environments.
  • Backup and disaster recovery - backup validation, disaster‑recovery architecture, failover testing, and recovery verification against defined RTO and RPO objectives.
  • Cloud security posture - cloud security posture management, runtime vulnerability detection, and configuration‑compliance monitoring, co‑owned with DevSecOps.
  • Data and streaming platform reliability - reliability engineering for Postgres, Redis/Valkey, Kafka, OpenSearch, and other critical stateful services.
  • Capacity, efficiency, and AI Ops - resource rightsizing, predictive alerting, autonomous triage, remediation automation, and self‑healing workflows.
  • Incident response and command - severity classification, incident command, executive communication, and blameless post‑incident review for the highest‑severity events.
  • On‑call and escalation - rotation design, operational readiness, escalation policy, and clean follow‑the‑sun handoffs across regions.
Responsibilities
  • Production reliability - own service‑level objectives, error budgets, and production reliability outcomes end to end, and represent reliability in architecture and scaling decisions.
  • Engineer observability - design and operate the observability stack so production services are fully instrumented and define the standards platform and application teams follow.
  • Govern SLOs and error budgets - define and manage service‑level objectives, run regular reviews with service owners, and drive corrective action when services consume error budgets unsafely.
  • Drive resilience - design and execute chaos experiments and validate that failure modes are covered by tested safeguards.
  • Lead incident response - define the incident process and serve as incident commander for the highest‑severity incidents, including security incidents within the coverage window.
  • Run on‑call and escalation - establish and manage rotations and escalation paths that provide continuous coverage with clean follow‑the‑sun handoffs.
  • Disaster recovery - own disaster‑recovery testing and failover validation against defined recovery objectives and turn exercise findings into architectural and operational improvements.
  • Cloud security posture - co‑own cloud security posture management, runtime vulnerability detection, and configuration‑compliance monitoring with DevSecOps.
  • Data, streaming, and AI Ops - own reliability of stateful and streaming services, capacity planning and rightsizing, and autonomous agents for triage, predictive alerting, remediation, and self‑healing.
  • Scale the team and broaden impact- mentor engineers at different seniority levels, set standards adopted across teams, and align Architecture, DevSecOps, Cloud Operations, and Product Development behind a shared reliability roadmap.
Requirements
  • 7+ years of production engineering experience with recent hands‑on reliability work.
  • Hands‑on, recent experience operating large‑scale, fault‑tolerant production systems on AWS or GCP.
  • Observability ownership - has instrumented production systems and governed service‑level objectives and error budgets, not only installed dashboards.
  • Resilience practice - has designed and executed failure experiments or disaster‑recovery exercises with real failover validation.
  • Incident command - has personally commanded serious SEV1/SEV2 incidents and driven root cause through to a systemic fix.
  • Demonstrated ownership of reliability outcomes with measurable results, such as availability, mean time to recovery, and error‑budget adherence.
  • Evidence of multi‑team technical leadership through standards, review, coaching, and mechanisms adopted beyond one service or team.
Preferred Qualifications
  • Proven production experience with cloud security posture management, runtime vulnerability detection, and workload protection across cloud and distributed environments.
  • Strong experience prioritizing risk using identity, workload, and exposure‑path context to focus remediation on issues that materially increase attack likelihood and operational impact.
  • Experience with autonomous remediation and self‑healing workflows powered by AIOps, including Amazon Bedrock Agent Core or equivalent agentic automation frameworks.
  • Hands‑on experience in capacity management, resource rightsizing, efficiency engineering, and practical cost optimization based on FinOps principles.
  • Experience with load‑balancing design and operations, including health‑based failover, global traffic management, and performance optimization for highly available services.
  • Experience with AI traffic management via an LLM gateway, including request routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and observability for multi‑model or multi‑provider environments.
  • Ability to connect networking, security, and reliability considerations into cohesive platform design decisions that improve resilience, performance, and operability.
Wage Transparency

The total compensation package includes base, bonus,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

IonQ • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

NightDragon Acquisition Corp. • Santa Clara (CA)

Hybrid
USD 152,000 - 228,000
Senior Staff Service Reliability and Operational Intelligence Engineer
Senior Staff Service Reliability and Operational Intelligence Engineer

Ionq • Santa Clara (CA)

On-site
USD 162,000 - 270,000
Staff Service Reliability and Operational Intelligence Engineer
Staff Service Reliability and Operational Intelligence Engineer

IonQ • Santa Clara (CA)

Hybrid
USD 180,000 - 280,000
Staff Service Reliability and Operational Intelligence Engineer
Staff Service Reliability and Operational Intelligence Engineer

NightDragon Acquisition Corp. • Santa Clara (CA)

Hybrid
USD 170,000 - 210,000
Home technology stipend
Medical, dental, vision coverage
401(k) matching
+1
Senior Distributed Systems Engineer
Senior Distributed Systems Engineer

Clutch Canada • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Staff Service Reliability and Operational Intelligence Engineer
Staff Service Reliability and Operational Intelligence Engineer

Physics World • Santa Clara (CA)

Hybrid
USD 170,000 - 250,000
Senior Staff Distributed Systems Engineer
Senior Staff Distributed Systems Engineer

Physics World • Santa Clara (CA)

On-site
USD 216,000 - 283,000
Staff DevOps Engineer
Staff DevOps Engineer

Physics World • Santa Clara (CA)

Hybrid
USD 150,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+5
Senior Distributed Systems Engineer
Senior Distributed Systems Engineer

Physics World • Santa Clara (CA)

On-site
USD 163,000 - 214,000