Reliability Engineer 3, Observability Specialist

Jobtailor

California (MO)

On-site

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior Observability Architect to own and drive the enterprise observability strategy across multiple platforms. You will define and govern enterprise-wide SLIs, SLOs, and reliability standards while architecting scalable telemetry-based solutions that span cloud, microservices, and Kubernetes environments.

You will advise executives and engineering leaders on production readiness, reliability improvements, and operational risk.

Qualifications

  • Bachelor's degree or equivalent work experience.
  • 5–7 years in IT service management, production support, or engineering roles related to observability and reliability.
  • Expertise in Observability Engineering or SRE/Reliability Engineering.
  • Strong knowledge of SLIs, SLOs, and Error Budgets.
  • Hands-on with APM, RUM, monitoring, logging, tracing, and telemetry frameworks.

Responsibilities

  • Own and drive the enterprise Observability Strategy across multiple platforms and tech domains.
  • Define and govern enterprise-wide SLIs, SLOs, Error Budgets, KPIs, and Reliability Standards.
  • Architect scalable observability solutions using telemetry, logging, metrics, and tracing.
  • Oversee governance frameworks, instrumentation standards, and alerting practices.
  • Advise leaders on reliability, production readiness, and resilience planning.
  • Lead executive and operational service health reporting on availability and reliability trends.
  • Drive continuous improvement through telemetry analysis and incident review.
  • Provide senior technical leadership and governance for observability practices.

Skills

Observability Engineering
Site Reliability Engineering
SLIs SLOs Error Budgets
APM RUM Telemetry
Dashboard & Alerts

Education

Bachelor's degree

Tools

Datadog
Dynatrace
Splunk
Grafana
Prometheus
New Relic
Elastic
OpenTelemetry

Job description

  • Own and drive the enterprise Observability Strategy across multiple platforms and technology domains
  • Define and govern enterprise-wide SLIs, SLOs, Error Budgets, KPIs, and Reliability Standards
  • Architect scalable observability solutions using telemetry, distributed tracing, logging, metrics, synthetic monitoring, and APM/RUM
  • Establish and oversee observability governance frameworks, instrumentation standards, telemetry policies, dashboard lifecycle management, alert governance, and monitoring best practices
  • Advise Product, Engineering, SRE, Infrastructure, and Operations leaders on technology strategy, production readiness, resilience planning, and operational risk management
  • Lead executive and operational service health reporting on availability, latency, customer impact, dependency performance, reliability trends, and SLO compliance
  • Drive continuous improvement through telemetry, incident, problem-management, and alert-effectiveness analysis
  • Provide senior technical leadership, mentorship, and standards governance for observability engineering practices across the enterprise
Requirements
  • Bachelor's degree, or equivalent work experience
  • Five to seven years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development
  • Expertise in Observability Engineering, Site Reliability Engineering (SRE), or Reliability Engineering
  • Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring
  • Ability to understand stakeholder needs and guide reliability requirements for large, complex multi-system products
  • Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks
  • Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry
  • Experience building, standardizing, and tuning operational dashboards and actionable alerts
  • Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes
  • Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements
  • Excellent stakeholder management, communication, and technical leadership skills
  • Must comply with U.S. Bank policies and procedures, including the Code of Ethics and Business Conduct and related workplace conduct and safety policies
  • Must be able to satisfy applicable employment eligibility verification and background-check requirements
Core Competencies

Demonstrates expertise in Observability Engineering and Site Reliability Engineering, with a strong focus on defining SLIs, SLOs, and Error Budgets. Capable of architecting scalable observability solutions and providing technical leadership across multiple technology domains.

Highest-signal resume keywords
  • Observability Engineering
  • Site Reliability Engineering (SRE)
  • SLIs, SLOs, Error Budgets
  • APM/RUM, Telemetry, Logging
  • Datadog, Dynatrace, Splunk
ATS Optimization Keywords
Hard Skills
  • Observability Strategy
  • Reliability Standards
  • Telemetry Frameworks
  • Incident Analysis
  • Performance Data Analysis
Soft Skills
  • Stakeholder Management
  • Technical Leadership
  • Communication
Industry Keywords
  • IT Service Management
  • Production Support
  • Product/Project Management
  • Cloud Platforms
  • Kubernetes
Tools & Technologies
  • Datadog
  • Dynatrace
  • Splunk
  • Grafana
  • Prometheus
  • New Relic
  • Elastic
  • OpenTelemetry
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Observability Engineer, Level 2
Reliability Observability Engineer, Level 2

Jobtailor • Colorado

On-site
USD 120,000 - 180,000
Head of Enterprise Monitoring and Observability
Head of Enterprise Monitoring and Observability

Jobtailor • Town of Bethlehem (NY)

On-site
USD 180,000 - 240,000
Engineering Manager, Developer Experience
Engineering Manager, Developer Experience

Jobtailor • California (MO)

On-site
USD 180,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Infosys • Richardson (TX)

On-site
USD 80,000 - 120,000
Distinguished Software Engineer – AI/ML Engineer
Distinguished Software Engineer – AI/ML Engineer

Jobtailor • Sunnyvale (CA)

On-site
USD 230,000 - 350,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • United States

On-site
USD 140,000 - 180,000
Observability Engineer
Observability Engineer

BCforward • Phoenix (AZ)

Hybrid
USD 120,000 - 140,000
SRE/Observability Engineer
SRE/Observability Engineer

BlueSky Resource Solutions • United States

Remote
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Stability Technology • United States

On-site
USD 120,000 - 170,000