Reliability Observability Engineer, Level 2

Jobtailor

Colorado

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor seeks a senior Observability Engineer to lead the strategy across critical customer journeys, aligning monitoring with business outcomes, reliability goals, and customer experience.

You will define SLIs/SLOs, governance, dashboards, and alerting, working with Product, Engineering, SRE, and Operations on production readiness and scalable telemetry. The role requires deep expertise in observability and cloud-native systems.

Qualifications

  • Bachelor's degree, or equivalent work experience.
  • 4–5 years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development.
  • Expertise in Observability Engineering, Site Reliability Engineering (SRE), or Reliability Engineering.
  • Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring.
  • Ability to understand stakeholder needs and guide development of reliability requirements for large, complex multi-system products.
  • Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks.
  • Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry.
  • Experience building, standardizing, and tuning operational dashboards and actionable alerts.
  • Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes.
  • Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements.
  • Excellent stakeholder management, communication, and technical leadership skills.
  • Applicants must be able to comply with U.S. Bank policies and procedures including the Code of Ethics and Business Conduct and related workplace conduct and safety policies

Responsibilities

  • Lead observability strategy across critical customer journeys, aligning monitoring capabilities with business outcomes, reliability goals, and customer experience.
  • Define, implement, and govern SLIs, SLOs, error budgets, and reliability metrics for enterprise applications and services.
  • Design and maintain scalable observability architectures, including telemetry instrumentation, monitoring frameworks, tagging standards, and alerting models.
  • Establish observability governance for dashboards, alerts, synthetic monitoring, telemetry standards, and monitoring-asset lifecycle management.
  • Partner with Product, Engineering, SRE, and Operations teams on production readiness, application instrumentation, and reliability measurement.
  • Develop and optimize service health dashboards and reporting covering availability, latency, customer impact, dependency performance, and SLO compliance.
  • Analyze telemetry data, incidents, alert history, problem records, and performance trends to identify gaps, reduce alert fatigue, and improve detection accuracy.
  • Provide technical leadership and mentorship on distributed tracing, logging, metrics, synthetic monitoring, APM, RUM, and alert governance best practices.

Skills

Observability Strategy
Telemetry Instrumentation
Monitoring Frameworks
Incident Analysis
Distributed Systems
Microservices
Cloud Platforms
Kubernetes
Service Health Dashboards
Reliability Metrics
Stakeholder Management
Communication
Technical Leadership

Education

Bachelor's degree or equivalent work experience

Tools

Datadog
Dynatrace
Splunk
Grafana
Prometheus
New Relic
Elastic
OpenTelemetry

Job description


  • Lead observability strategy across critical customer journeys, aligning monitoring capabilities with business outcomes, reliability goals, and customer experience

  • Define, implement, and govern SLIs, SLOs, error budgets, and reliability metrics for enterprise applications and services

  • Design and maintain scalable observability architectures, including telemetry instrumentation, monitoring frameworks, tagging standards, and alerting models

  • Establish observability governance for dashboards, alerts, synthetic monitoring, telemetry standards, and monitoring-asset lifecycle management

  • Partner with Product, Engineering, SRE, and Operations teams on production readiness, application instrumentation, and reliability measurement

  • Develop and optimize service health dashboards and reporting covering availability, latency, customer impact, dependency performance, and SLO compliance

  • Analyze telemetry data, incidents, alert history, problem records, and performance trends to identify gaps, reduce alert fatigue, and improve detection accuracy

  • Provide technical leadership and mentorship on distributed tracing, logging, metrics, synthetic monitoring, APM, RUM, and alert governance best practices


Requirements


  • Bachelor's degree, or equivalent work experience

  • Four to five years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development

  • Expertise in Observability Engineering, Site Reliability Engineering (SRE), or Reliability Engineering

  • Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring

  • Ability to understand stakeholder needs and guide development of reliability requirements for large, complex multi-system products

  • Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks

  • Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry

  • Experience building, standardizing, and tuning operational dashboards and actionable alerts

  • Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes

  • Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements

  • Excellent stakeholder management, communication, and technical leadership skills

  • Applicants must be able to comply with U.S. Bank policies and procedures including the Code of Ethics and Business Conduct and related workplace conduct and safety policies


Core Competencies

Demonstrates expertise in Observability Engineering and Site Reliability Engineering, focusing on defining and implementing SLIs, SLOs, and reliability metrics. Proficient in leveraging telemetry data and monitoring frameworks to enhance service health and customer experience.


Highest-signal resume keywords


  • Observability Engineering

  • Site Reliability Engineering (SRE)

  • SLIs, SLOs, Error Budgets

  • APM, RUM, Monitoring Frameworks

  • Datadog, Dynatrace, Splunk


Hard Skills


  • Observability Strategy

  • Telemetry Instrumentation

  • Monitoring Frameworks

  • Incident Analysis

  • Distributed Systems

  • Microservices

  • Cloud Platforms

  • Kubernetes

  • Service Health Dashboards

  • Reliability Metrics


Soft Skills


  • Stakeholder Management

  • Communication

  • Technical Leadership


Industry Keywords


  • IT Service Management

  • Production Support

  • Product Management

  • Application Development

  • Risk Analysis


Tools & Technologies


  • Datadog

  • Dynatrace

  • Splunk

  • Grafana

  • Prometheus

  • New Relic

  • Elastic

  • OpenTelemetry

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer 3, Observability Specialist
Reliability Engineer 3, Observability Specialist

Jobtailor • California (MO)

On-site
USD 140,000 - 190,000
Head of Enterprise Monitoring and Observability
Head of Enterprise Monitoring and Observability

Jobtailor • Town of Bethlehem (NY)

On-site
USD 180,000 - 240,000
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • United States

On-site
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Infosys • Richardson (TX)

On-site
USD 80,000 - 120,000
Engineering Manager, Developer Experience
Engineering Manager, Developer Experience

Jobtailor • California (MO)

On-site
USD 180,000 - 250,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Senior Reliability Engineer
Senior Reliability Engineer

Jobtailor • Menomonee Falls (WI)

On-site
USD 120,000 - 160,000