Senior Observability Engineer, Unified Platform

NCR Corporation

Atlanta (GA)

On-site

USD 160,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NCR Voyix Corporation seeks a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and maturation of a unified observability platform across NCR Voyix Restaurants, Retail, and Payments environments. The role requires 10+ years in SRE/Platform/DevOps with a track record in large-scale enterprise observability.

The candidate will collaborate with Product, Infrastructure, Security, Operations, and Leadership to set standards and roadmaps,

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • 10+ years of experience in Site Reliability Engineering, Cloud Engineering, Platform Engineering, DevOps, or related disciplines.
  • Proven design and operation of observability platforms in large-scale enterprise environments.
  • Deep expertise with Kubernetes platforms, including AKS and GKE.
  • Strong experience with Azure and Google Cloud Platform services and architectures.
  • Hands-on experience with enterprise observability tools such as Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic, or similar platforms.
  • Advanced knowledge of monitoring, logging, telemetry collection, distributed tracing, and observability engineering principles.
  • Experience defining and operationalizing SLIs, SLOs, error budgets, reliability metrics, and service health frameworks.
  • Strong automation and Infrastructure as Code using Terraform and related tools.
  • Proficiency developing automation solutions using Python, Go, PowerShell, or similar languages.
  • Experience integrating observability solutions into CI/CD pipelines and modern DevOps workflows.
  • Proven ability to influence technical direction and collaborate with stakeholders across Engineering, Product, Infrastructure, Security, and Operations.
  • Strong communication, leadership, and stakeholder management skills.

Responsibilities

  • Lead the architecture, design, implementation, and continuous improvement of enterprise observability solutions across Azure, Google Cloud Platform (GCP), Kubernetes, and hybrid environments.
  • Establish and drive enterprise observability standards for monitoring, logging, distributed tracing, telemetry, and operational analytics.
  • Develop and maintain dashboards that provide real-time visibility into infrastructure, applications, platform health, customer experience, and business transactions.
  • Define, evangelize, and implement reliability frameworks including SLIs, SLOs, error budgets, operational KPIs, and service health metrics.
  • Partner cross-functionally with Engineering, Infrastructure, Security, Product, and Operations teams to identify reliability risks.
  • Lead efforts to improve incident prevention, detection, response, and recovery through intelligent alerting, automation, event correlation, and observability best practices.
  • Integrate observability capabilities with ServiceNow, CI/CD pipelines, automation frameworks, and enterprise workflows.
  • Influence technical strategy and roadmap decisions related to reliability engineering and platform observability.
  • Support and drive enterprise initiatives involving AI-driven observability, predictive analytics, and anomaly detection.
  • Mentor engineers and serve as SME for observability and cloud-native operations.
  • Establish governance and best practices across product and engineering teams to ensure enterprise-wide standards.

Skills

Kubernetes
AKS
GKE
Azure
Google Cloud
Grafana
Datadog
Prometheus
OpenTelemetry
Dynatrace
New Relic
Terraform
Python
Go
PowerShell
CI/CD
Automation
ServiceNow

Education

Bachelor's degree in Computer Science / IT / Engineering

Tools

Grafana
Datadog
Prometheus
OpenTelemetry
Dynatrace
New Relic

Job description

NCR Voyix Corporation seeks a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and maturation of a unified observability platform across NCR Voyix Restaurants, Retail, and Payments environments. The role requires 10+ years in SRE/Platform/DevOps with a track record in large-scale enterprise observability.

The candidate will collaborate with Product, Infrastructure, Security, Operations, and Leadership to set standards and roadmaps,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer – Unified Observability
Senior Site Reliability Engineer – Unified Observability

NCR Corporation • Atlanta (GA)

On-site
USD 160,000 - 210,000
Senior Observability Platform Engineering Lead
Senior Observability Platform Engineering Lead

CVS Health • Town of Florida (NY)

Hybrid
USD 118,000 - 237,000
Medical, dental, vision coverage
Paid time off
Retirement options
Senior Observability Platform Engineer: Scale & Reliability
Senior Observability Platform Engineer: Scale & Reliability

CVS Health • Maine

Hybrid
USD 83,000 - 222,000
AI-Driven Observability Engineer for Automation & SRE
AI-Driven Observability Engineer for Automation & SRE

NCR Corporation • Atlanta (GA), Northern (KY)

Hybrid
USD 120,000 - 170,000
AI-Driven Observability Engineer
AI-Driven Observability Engineer

NCR Voyix • Atlanta (GA)

On-site
USD 120,000 - 150,000
Senior SRE & Platform Engineer — Observability & Automation
Senior SRE & Platform Engineer — Observability & Automation

Techunting • United States

On-site
USD 120,000 - 150,000
AI-Driven Observability Engineer
AI-Driven Observability Engineer

ncr • Atlanta (GA)

On-site
USD 110,000 - 170,000
Senior SRE, Observability for Retail Platforms (Hybrid NYC)
Senior SRE, Observability for Retail Platforms (Hybrid NYC)

Tech Mirrors • New York (NY)

Hybrid
USD 140,000 - 190,000
Staff SRE: Observability & Reliability Platform Lead
Staff SRE: Observability & Reliability Platform Lead

Koitecc Solutions • Scottsdale (AZ)

On-site
USD 118,000 - 261,000
Medical, dental, vision benefits
Paid time off
Retirement plan and equity options
+1
Staff SRE: Architect of Observability & Reliability Platform
Staff SRE: Architect of Observability & Reliability Platform

CVS Health • Richardson (TX)

On-site
USD 118,000 - 261,000