Senior Observability & Reliability Engineer

U.S. Bank

Cupertino (CA)

On-site

USD 124,000 - 146,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare
Life insurance
Disability leave
401(k) retirement plan
Paid vacation
Paid holidays

Job summary

U.S. Bank is seeking a senior Reliability Engineer specializing in observability to translate customer journeys into measurable reliability objectives. You will lead the creation and governance of SLIs, SLOs, dashboards, and alerts across complex enterprise apps, ensuring instrumented systems and data-driven reliability decisions.

Collaboration with product, SRE, and operations is essential. The role emphasizes incident analysis, reducing alert fatigue, and mentoring teams on tracing, metrics,

Qualifications

  • Six to eight years of relevant work experience in IT and reliability.
  • Strong knowledge of SLIs, SLOs, and error budgets.
  • Experience building dashboards and governance for observability assets.

Responsibilities

  • Lead the definition, documentation, implementation, and continuous improvement of Observability across Critical Customer Journeys, ensuring alignment between Observability Strategy, business outcomes, and reliability objectives.
  • Design, implement, and govern SLIs, SLOs, Error Budgets, and reliability metrics for enterprise applications and services.
  • Establish and maintain Observability Governance Frameworks for Dashboards, Alerts, Synthetic Monitoring, Telemetry Standards, and lifecycle management of observability assets.
  • Translate business and technical requirements into scalable Observability Architectures, including Instrumentation Standards, Monitoring Strategies, Tagging Frameworks, and Alerting Models.
  • Partner with Product Owners, Application Engineering, Site Reliability Engineering (SRE), and Operations Teams to ensure applications are production-ready and fully instrumented for reliability measurement.
  • Develop and maintain executive and operational Service Health Dashboards that provide insights into Availability, Latency, Customer Impact, Dependency Performance, and SLO Compliance.
  • Analyze Telemetry Data, Incident Trends, Problem Records, and Alert Performance to identify observability gaps, reduce alert fatigue, and improve detection accuracy.
  • Provide technical leadership and mentorship on Distributed Tracing, Logging, Metrics Collection, Synthetic Testing, Application Performance Monitoring, Real User Monitoring, Monitoring Design Patterns, and Alert Governance Best Practices.
  • Lead the definition, documentation, and ongoing refinement of critical user journeys in partnership with product owners, engineering teams, SRE, operations, and business stakeholders to ensure observability practices are aligned to customer experience, business outcomes, and operational risk.
  • Define, document, and govern appropriate service-level indicators and service-level objectives for applications and key capabilities, including availability, latency, error rate, throughput, dependency health, and other measurements that reflect meaningful customer and business impact.
  • Establish and maintain a best-practice process for identifying, approving, implementing, reviewing, and retiring user journeys, SLIs, SLOs, dashboards, monitors, synthetic tests, alerts, and related observability artifacts.
  • Translate product and engineering requirements into actionable observability designs that specify telemetry needs, measurement methods, tagging standards, dashboard requirements, alerting thresholds, ownership, evidence expectations, and operational runbook linkages.
  • Partner with product and engineering teams during design, build, release, and production-readiness activities to ensure applications are instrumented to validate critical customer journeys, measure reliability outcomes, and support effective incident detection and triage.
  • Develop, maintain, and continuously improve dashboards and reporting that communicate service health, SLO performance, error-budget posture, customer impact, dependency performance, alert effectiveness, and trends to technical teams and leadership stakeholders.
  • Analyze telemetry, incidents, problem records, alert history, customer-impacting events, and SLO performance trends to identify observability gaps, reduce alert noise, improve detection accuracy, and recommend reliability improvements.
  • Provide senior-level guidance, coaching, and standards interpretation to engineering, SRE, and operations teams on observability design patterns, SLI/SLO selection, customer journey monitoring, synthetic monitoring, logging, tracing, metrics, and alert governance.
  • Maintain an authoritative inventory of observability assets, including user journeys, SLIs, SLOs, dashboards, monitors, alerts, synthetic tests, ownership assignments, review cadence, and evidence of ongoing compliance with approved observability standards.

Skills

Observability Engineering
SRE/Reliability
SLIs/SLOs
Distributed Tracing
Monitoring & Alerts
Telemetry/Logging
Stakeholder Communication
Mentorship
Incident Analysis

Education

Bachelor's degree

Tools

Datadog
Dynatrace
Splunk
Grafana
Prometheus
New Relic
Elastic
OpenTelemetry

Job description

U.S. Bank is seeking a senior Reliability Engineer specializing in observability to translate customer journeys into measurable reliability objectives. You will lead the creation and governance of SLIs, SLOs, dashboards, and alerts across complex enterprise apps, ensuring instrumented systems and data-driven reliability decisions.

Collaboration with product, SRE, and operations is essential. The role emphasizes incident analysis, reducing alert fatigue, and mentoring teams on tracing, metrics,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

Us Bank • Chicago (IL)

On-site
USD 124,000 - 146,000
Healthcare
Retirement plan
Paid vacation
+2
Senior Observability Engineer: Reliability Focus
Senior Observability Engineer: Reliability Focus

Relha LLC • Town of Brookfield (WI)

Hybrid
USD 98,000 - 116,000
Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

Us Bank • Cincinnati (OH)

On-site
USD 98,000 - 116,000
Healthcare
401(k) plan
Paid vacation
+1
Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

U.S. Bank • Town of Brookfield (WI), Northern (KY)

Hybrid
USD 98,000 - 116,000
Healthcare (medical, dental, vision)
Retirement plan (401(k))
Paid vacation
+2
Senior Observability & Reliability Engineer (SRE)
Senior Observability & Reliability Engineer (SRE)

U.S. Bank • Cupertino (CA)

On-site
USD 98,000 - 116,000
Healthcare
Retirement plan
Paid time off
+1
Lead Observability Engineer (SRE) — Reliability & Dashboards
Lead Observability Engineer (SRE) — Reliability & Dashboards

Relha LLC • Atlanta (GA), Northern (KY)

Hybrid
USD 86,000 - 102,000
401(k) retirement plan
Paid vacation
Up to 11 paid holidays
+2
Senior Observability & Reliability Engineer (SLI/SLO)
Senior Observability & Reliability Engineer (SLI/SLO)

Us Bank • Town of Brookfield (WI)

On-site
USD 98,000 - 116,000
Healthcare
401(k)
Paid vacation
+3
Lead Observability Engineer - Reliability & Metrics
Lead Observability Engineer - Reliability & Metrics

U.S. Bank • Atlanta (GA)

On-site
USD 86,000 - 102,000
Healthcare (medical, dental, vision)
401(k) and employer-funded retirement
Paid vacation and holidays
Reliability Engineer 3 (Observability Specialist)
Reliability Engineer 3 (Observability Specialist)

U.S. Bank • Town of Brookfield (WI), Northern (KY)

Hybrid
USD 98,000 - 116,000
Healthcare (medical, dental, vision)
Retirement plan (401(k))
Paid vacation
+2
Reliability Engineer 3 (Observability Specialist)
Reliability Engineer 3 (Observability Specialist)

U.S. Bank • Cupertino (CA)

On-site
USD 98,000 - 116,000
Healthcare
Retirement plan
Paid time off
+1