Senior Staff Reliability & Observability Architect

Ionq

Santa Clara (CA)

On-site

USD 162,000 - 270,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IonQ is seeking a Senior Staff Service Reliability and Operational Intelligence Engineer to shape reliability across regions and services. You will design and operate observability platforms, govern SLO programs, lead high-severity incident response, and drive self-healing through AI Ops workflows.

You will own the reliability governance model for production services, align with business impact, and mentor teams while delivering scalable, secure, and resilient cloud-native platforms on AWS/GCP.

Qualifications

  • 12+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
  • Recent experience designing and operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Deep understanding of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes.
  • Demonstrated ownership of observability architecture, including instrumentation of production systems and governance of metrics, logs, traces, SLIs, SLOs, and error budgets.

Responsibilities

  • Own the technical strategy and multi-year roadmap for operational excellence and production readiness across development, pre-production, and production environments.
  • Define and govern the New Service Introduction framework, including mandatory architecture, security, resilience, capacity, observability, supportability, and release-readiness reviews before services enter production.
  • Establish organization-wide service ownership standards covering service catalog records, accountable owners, dependency maps, runbooks, support models, escalation paths, recovery objectives, and on-call readiness.
  • Lead the architecture and evolution of the shared observability platform, establishing consistent standards for logs, metrics, distributed traces, and profiles across production systems.

Skills

Production engineering
SRE
Observability
Kubernetes
AWS
GCP
Incident response
Python
Go
IaC

Tools

Python
Go
Terraform
CI/CD

Job description

IonQ is seeking a Senior Staff Service Reliability and Operational Intelligence Engineer to shape reliability across regions and services. You will design and operate observability platforms, govern SLO programs, lead high-severity incident response, and drive self-healing through AI Ops workflows.

You will own the reliability governance model for production services, align with business impact, and mentor teams while delivering scalable, secure, and resilient cloud-native platforms on AWS/GCP.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff Reliability & AI Ops Engineer
Senior Staff Reliability & AI Ops Engineer

IonQ, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 188,000 - 270,000
Medical plan
Dental plan
Vision plan
+5
Senior Staff Service Reliability and Operational Intelligence Engineer
Senior Staff Service Reliability and Operational Intelligence Engineer

Ionq • Santa Clara (CA)

On-site
USD 162,000 - 270,000
Senior Staff Distributed Systems Engineer — Go/Rust Backend
Senior Staff Distributed Systems Engineer — Go/Rust Backend

IonQ • Santa Clara (CA)

On-site
USD 216,000 - 283,000
Senior Site Reliability Engineer — Scale & Observability
Senior Site Reliability Engineer — Scale & Observability

Pivotal Health • New York (NY)

Hybrid
USD 230,000 - 260,000
Equity
Health, dental, vision
401(k)
+2
Senior Staff Service Reliability and Operational Intelligence Engineer New Santa Clara, California, United States
Senior Staff Service Reliability and Operational Intelligence Engineer New Santa Clara, California, United States

IonQ, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 188,000 - 270,000
Medical plan
Dental plan
Vision plan
+5
Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior Observability & Automation Architect
Senior Observability & Automation Architect

Q2 India • Austin (TX)

Hybrid
USD 170,000 - 210,000
Health & Wellness Benefits
Hybrid Work Opportunities
Generous parental leave
+3
Senior Site Reliability Engineer: Cloud, CI/CD & Observability
Senior Site Reliability Engineer: Cloud, CI/CD & Observability

Castleton Commodities International, LLC • United States

On-site
USD 160,000 - 260,000
Medical & Dental
Pension Plan
Tuition assistance
+2
Staff Infra Engineer - Reliability & Observability (Remote)
Staff Infra Engineer - Reliability & Observability (Remote)

Wisdom • United States

Remote
USD 120,000 - 160,000
Senior SRE: AI-Driven Infra & Reliability
Senior SRE: AI-Driven Infra & Reliability

Jobless • Ann Arbor (MI)

Hybrid
USD 180,000 - 240,000
Health Care Coverage
Life Insurance
Health Savings Account
+3