Observability & Reliability Engineer (SRE)

Jobtailor

Colorado

On-site

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor seeks a senior Observability Engineer to lead the strategy across critical customer journeys, aligning monitoring with business outcomes, reliability goals, and customer experience.

You will define SLIs/SLOs, governance, dashboards, and alerting, working with Product, Engineering, SRE, and Operations on production readiness and scalable telemetry. The role requires deep expertise in observability and cloud-native systems.

Qualifications

  • Bachelor's degree, or equivalent work experience.
  • 4–5 years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development.
  • Expertise in Observability Engineering, Site Reliability Engineering (SRE), or Reliability Engineering.
  • Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring.
  • Ability to understand stakeholder needs and guide development of reliability requirements for large, complex multi-system products.
  • Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks.
  • Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry.
  • Experience building, standardizing, and tuning operational dashboards and actionable alerts.
  • Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes.
  • Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements.
  • Excellent stakeholder management, communication, and technical leadership skills.
  • Applicants must be able to comply with U.S. Bank policies and procedures including the Code of Ethics and Business Conduct and related workplace conduct and safety policies

Responsibilities

  • Lead observability strategy across critical customer journeys, aligning monitoring capabilities with business outcomes, reliability goals, and customer experience.
  • Define, implement, and govern SLIs, SLOs, error budgets, and reliability metrics for enterprise applications and services.
  • Design and maintain scalable observability architectures, including telemetry instrumentation, monitoring frameworks, tagging standards, and alerting models.
  • Establish observability governance for dashboards, alerts, synthetic monitoring, telemetry standards, and monitoring-asset lifecycle management.
  • Partner with Product, Engineering, SRE, and Operations teams on production readiness, application instrumentation, and reliability measurement.
  • Develop and optimize service health dashboards and reporting covering availability, latency, customer impact, dependency performance, and SLO compliance.
  • Analyze telemetry data, incidents, alert history, problem records, and performance trends to identify gaps, reduce alert fatigue, and improve detection accuracy.
  • Provide technical leadership and mentorship on distributed tracing, logging, metrics, synthetic monitoring, APM, RUM, and alert governance best practices.

Skills

Observability Strategy
Telemetry Instrumentation
Monitoring Frameworks
Incident Analysis
Distributed Systems
Microservices
Cloud Platforms
Kubernetes
Service Health Dashboards
Reliability Metrics
Stakeholder Management
Communication
Technical Leadership

Education

Bachelor's degree or equivalent work experience

Tools

Datadog
Dynatrace
Splunk
Grafana
Prometheus
New Relic
Elastic
OpenTelemetry

Job description

Jobtailor seeks a senior Observability Engineer to lead the strategy across critical customer journeys, aligning monitoring with business outcomes, reliability goals, and customer experience.

You will define SLIs/SLOs, governance, dashboards, and alerting, working with Product, Engineering, SRE, and Operations on production readiness and scalable telemetry. The role requires deep expertise in observability and cloud-native systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Enterprise Observability & Reliability Lead
Enterprise Observability & Reliability Lead

Jobtailor • California (MO)

On-site
USD 140,000 - 190,000
Senior SRE Leader: Reliability & Observability
Senior SRE Leader: Reliability & Observability

Expedite Talent Solutions • United States

Hybrid
USD 130,000 - 160,000
Senior Observability & Reliability Engineer (SRE)
Senior Observability & Reliability Engineer (SRE)

U.S. Bank • Cupertino (CA)

On-site
USD 98,000 - 116,000
Healthcare
Retirement plan
Paid time off
+1
Lead Observability Engineer (SRE) — Reliability & Dashboards
Lead Observability Engineer (SRE) — Reliability & Dashboards

Relha LLC • Atlanta (GA), Northern (KY)

Hybrid
USD 86,000 - 102,000
401(k) retirement plan
Paid vacation
Up to 11 paid holidays
+2
Observability & SRE Engineer — Cloud, Kubernetes, & Automation
Observability & SRE Engineer — Cloud, Kubernetes, & Automation

Ontrac Solutions • Arizona

Hybrid
USD 140,000 - 190,000
Verification cost reimbursement
SRE Architect — Reliability & Observability Leader
SRE Architect — Reliability & Observability Leader

Vkore Solutions • City of Albany (NY)

On-site
USD 140,000 - 180,000
Technical Lead SRE — Reliability & Observability
Technical Lead SRE — Reliability & Observability

LSEG • Raleigh (NC)

On-site
USD 140,000 - 190,000
Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

Us Bank • Chicago (IL)

On-site
USD 124,000 - 146,000
Healthcare
Retirement plan
Paid vacation
+2
Senior SRE & Platform Engineer — Observability & Automation
Senior SRE & Platform Engineer — Observability & Automation

Techunting • United States

On-site
USD 120,000 - 150,000
Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

U.S. Bank • Cupertino (CA)

On-site
USD 124,000 - 146,000
Healthcare
Life insurance
Disability leave
+3