Staff Observability Platform Engineer (SRE)

CVS Health

Richardson (TX)

On-site

USD 118,450 - 236,900

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision coverage
Paid time off
Retirement savings options
Wellness programs

Job summary

Hispanic Alliance for Career Enhancement in Richardson, Texas, is seeking a Health Systems Engineer. This role requires over 10 years of software engineering experience, focusing on developing performance metrics and monitoring solutions. The ideal candidate will have proficiency in Java or Python, experience with cloud platforms, and strong analytical skills. Competitive salary range is between $118,450 and $236,900 with benefits that include medical, dental, and wellness programs.

Qualifications

  • 10+ years of experience in Software Engineering, Platform Engineering, or SRE.
  • 7+ years of experience with observability practices like alerting and incident management.
  • 7+ years building production-grade backend services in Java or Python.

Responsibilities

  • Define and maintain key performance metrics to measure system reliability.
  • Collaborate with development teams to manage error budgets effectively.
  • Design and implement monitoring solutions for system health visibility.

Skills

Software Engineering
Platform Engineering
Java
Python
Observability practices
Cloud platforms (AWS, GCP, Azure)
Docker
Kubernetes

Education

Bachelor's degree or equivalent experience

Tools

OpenTelemetry
Grafana
Prometheus
PostgreSQL
MySQL

Job description

POSITION SUMMARY

We're building a world of health around every individual – shaping a more connected, convenient, and compassionate health experience. At CVS Health®, you'll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable, and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family, and one community at a time.

Responsibilities
  • Metrics Development: Define, implement, and maintain key performance metrics, SLOs, and SLIs to measure system reliability and performance. Ensure alignment with business objectives and operational goals.
  • Error Budgets: Manage error budgets effectively, collaborating with development teams to balance reliability and feature delivery. Analyze incidents and outages to inform adjustments to error budgets.
  • Monitoring & Observability: Design and implement comprehensive monitoring solutions to provide real‑time visibility into system health. Utilize tools such as Prometheus, Grafana, Loki, Tempo, and other observability platforms to create dashboards and alerts.
  • Cloud Infrastructure Scaling: Architect, design, and implement scalable cloud infrastructure capable of supporting multiple business applications, ensuring reliability, performance, and future growth.
  • Quality Gates Automation: Develop and implement automated quality gates that ensure all releases meet defined reliability and performance standards. Lead the release DevOps team to integrate these gates into the CI/CD pipeline.
  • Incident Management: Assist in incident response efforts by providing insights from metrics and monitoring tools. Conduct post‑mortem analyses to identify root causes and recommend preventive measures.
Required Qualifications
  • 10+ years of experience in Software Engineering, Platform Engineering, or SRE.
  • 7+ years of experience with observability practices, including SLIs/SLOs/SLAs, alerting, and incident management.
  • 7+ years building production‑grade backend services in Java or Python.
  • 7+ years implementing and operating OpenTelemetry, including OTLP, semantic conventions, and instrumentation patterns.
  • 7+ years with cloud‑native and containerized platforms (Docker, Kubernetes, Argo CD).
  • 7+ years working with public cloud platforms (AWS, GCP, or Azure).
  • 5+ years designing and scaling distributed, high‑volume data pipelines.
  • 5+ years working with Grafana OSS or comparable observability backends (e.g., Grafana, Loki, Tempo, Prometheus).
  • 5+ years with relational databases (PostgreSQL, MySQL).
Preferred Qualifications
  • Excellent analytical skills and the ability to communicate complex technical concepts to non‑technical stakeholders.
  • Experience with service meshes and networking technologies such as Envoy and Istio.
  • Experience integrating or operating commercial observability platforms (Splunk, AppDynamics, etc.).
  • Experience with streaming and data platforms such as Kafka, Pulsar, or similar technologies.
  • Familiarity with time‑series, NoSQL, or analytical databases (ClickHouse, Bigtable, Cassandra, etc.).
  • Experience with Infrastructure as Code tools such as Terraform or CloudFormation.
  • Experience with cost optimization and capacity planning for large‑scale cloud infra.
  • Experience with chaos engineering, resiliency testing, or fault injection.
  • Background in security‑aware platform design, including secure service‑to‑service communication.
  • Experience mentoring senior engineers and influencing platform standards across organizations.
  • Strong operational experience supporting 24x7 production systems, including on‑call responsibilities.
  • Knowledge of security best practices in cloud environments.
Education

Bachelor's degree or equivalent experience (HS diploma + 4 years relevant experience)

Pay Range

The typical pay range for this role is: $118,450.00 – $236,900.00

Benefits

We take pride in offering a comprehensive and competitive mix of pay and benefits that reflects our commitment to our colleagues and their families. This full‑time position is eligible for a comprehensive benefits package designed to support the physical, emotional, and financial well‑being of colleagues and their families. The benefits for this position include medical, dental, and vision coverage, paid time off, retirement savings options, wellness programs, and other resources, based on eligibility.

EEO Statement

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Observability Platform Engineer (SRE)
Staff Observability Platform Engineer (SRE)

Koitecc Solutions • Richardson (TX)

Hybrid
USD 118,000 - 237,000
Staff Observability Platform Engineer (SRE)
Staff Observability Platform Engineer (SRE)

Koitecc Solutions • Scottsdale (AZ)

On-site
USD 118,000 - 237,000
Great benefits
Senior Engineer - Observability Platform
Senior Engineer - Observability Platform

CVS Health • Nevada (IA)

Hybrid
USD 83,000 - 222,000
Senior Engineer - Observability Platform
Senior Engineer - Observability Platform

CVS Health • Virginia (IL)

Hybrid
USD 83,000 - 222,000
Senior Engineer - Observability Platform
Senior Engineer - Observability Platform

CVS Health • New York (NY)

Hybrid
USD 83,000 - 222,000
Senior Engineer - Observability Platform
Senior Engineer - Observability Platform

CVS Health • Tennessee

Hybrid
USD 83,000 - 222,000
Medical, dental, and vision coverage
Paid time off
Retirement savings options
+1
Senior Engineer - Observability Platform
Senior Engineer - Observability Platform

CVS Health • Wisconsin

Hybrid
USD 83,000 - 222,000
Senior Engineer - Observability Platform
Senior Engineer - Observability Platform

CVS Health • Arkansas

Hybrid
USD 83,000 - 222,000
Senior Engineer - Observability Platform
Senior Engineer - Observability Platform

CVS Health • Maryland

Hybrid
USD 83,000 - 222,000
Medical, dental, and vision coverage
Paid time off
Retirement savings options
+1
Senior Engineer - Observability Platform
Senior Engineer - Observability Platform

CVS Health • North Dakota

Hybrid
USD 83,000 - 222,000