Site Reliability Engineer Shift Manager

NCR Voyix

Chennai District

On-site

INR 1,200,000 - 2,400,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NCR Voyix is seeking a Global Command Center Site Reliability Engineer to ensure the availability, reliability, performance, and operational stability of business-critical applications and services across hybrid cloud environments. The role collaborates with application, infrastructure, and cloud teams to prevent outages and drive improvements.

The position emphasizes proactive monitoring, incident response, automation, and governance, with responsibilities spanning ORR participation, alert

Qualifications

  • 2+ years of experience in Site Reliability Engineering, Production Support, Systems Engineering, DevOps, or NOC.
  • Experience with incident, problem, change, and service level management.
  • Experience with ServiceNow or similar ITSM platforms.

Responsibilities

  • Monitor enterprise applications, infrastructure, cloud platforms, and customer-facing services to ensure availability and performance.
  • Develop dashboards, alerts, and health monitoring solutions; reduce alert fatigue through tuning.
  • Automate operational processes and tooling; support infrastructure-as-code initiatives.

Skills

Site Reliability Engineering
Production Support
Systems Engineering
DevOps
NOC
Command Center Operations

Tools

AppDynamics
Datadog
Dynatrace
New Relic
Splunk
Azure Monitor
ServiceNow

Job description

Job Description:

NCR Voyix Corporation (NYSE: VYX) is a global platform-powered leader in unified commerce for shopping and dining. Combining a flexible, intelligent platform with end-to-end payments capabilities and services developed through its deep industry experience, NCR Voyix empowers retailers and restaurants to accelerate new possibilities for their operations, experiences and business outcomes. NCR Voyix is headquartered in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.

Global Command Center (GCC) Site Reliability Engineer
Position Summary

The Global Command Center (GCC) Site Reliability Engineer (I) is responsible for ensuring the availability, reliability, performance, and operational stability of business-critical applications, infrastructure, and customer-facing services. This role serves as a key member of the Global Command Center, providing proactive monitoring, incident response, problem management, automation, and operational excellence across enterprise environments.

The GCC works closely with application teams, infrastructure teams, cloud engineering, network operations, vendors, and business stakeholders to prevent outages, reduce operational risk, and drive continuous service improvement. The role participates in major incident management, root cause analysis, operational readiness reviews, and reliability engineering initiatives.

Key Responsibilities
Reliability & Service Availability
  • Monitor enterprise applications, infrastructure, cloud platforms, and customer-facing services.
  • Ensure platform availability, performance, and service health within established SLAs and SLOs.
  • Identify service degradation trends and proactively address reliability risks.
  • Drive operational improvements to reduce incidents and improve system resiliency.
  • Establish and track reliability metrics, KPIs, and service health indicators.
  • Develop preventative measures to reduce future service disruptions.
Observability & Monitoring
  • Configure and maintain monitoring platforms including:
    • AppDynamics
    • Dynatrace
    • Datadog
    • Splunk
    • Azure Monitor
    • Google Cloud Operations
    • ServiceNow Event Management
  • Develop dashboards, alerts, and health monitoring solutions.
  • Reduce alert fatigue through alert tuning and optimization.
Automation & Engineering
  • Automate operational processes and repetitive tasks.
  • Develop scripts and tooling using:
    • PowerShell
    • Python
    • Bash
    • APIs
  • Improve operational efficiency through self-healing and automated remediation capabilities.
  • Support infrastructure-as-code and reliability engineering initiatives.
Operational Readiness
  • Participate in Operational Readiness Reviews (ORR).
  • Validate monitoring, alerting, runbooks, and support procedures prior to production go-live.
  • Ensure escalation paths and support models are documented and operational.
  • Review application deployments for supportability and operational risks.
Cloud & Infrastructure Support
  • Support hybrid environments across:
    • Azure
    • Google Cloud Platform (GCP)
    • AWS
    • VMware
  • Analyze application, network, database, and infrastructure performance issues.
  • Work with engineering teams to optimize platform stability and scalability.
Governance & Reporting
  • Produce incident reports, service health updates, operational reviews, and executive summaries.
  • Maintain operational documentation, runbooks, and knowledge articles.
  • Track service performance metrics and reliability improvements.
  • Support audit and compliance initiatives as required.
Required Qualifications
  • 2+ years of experience in:
    • Site Reliability Engineering
    • Production Support
    • Systems Engineering
    • DevOps
    • Network Operations Center (NOC)
    • Command Center Operations
  • Experience supporting mission-critical production environments.
  • Strong understanding of:
    • Incident Management
    • Problem Management
    • Change Management
    • Service Level Management
  • Experience with ServiceNow or similar ITSM platforms.
Technical Skills
Operating Systems
  • Windows Server
  • Linux/Unix
Cloud Platforms
  • Microsoft Azure
  • Google Cloud Platform (GCP)
  • Amazon Web Services (AWS)
Monitoring & Observability
  • AppDynamics
  • Datadog
  • Dynatrace
  • New Relic
  • Splunk
  • Azure Monitor
  • ServiceNow Event Management
Preferred Qualifications
  • Experience working in a Global Command Center environment.
  • AWS, Azure, or GCP certifications.
  • Experience supporting retail, hospitality, payments, or enterprise SaaS platforms.
  • Experience with CI/CD pipelines and DevOps practices.
  • Knowledge of SRE concepts including:
    • SLI/SLO/SLA management
    • Error budgets
    • Chaos testing
    • Resiliency engineering
Key Competencies
  • Critical Incident Leadership
  • Technical Troubleshooting
  • Problem Solving
  • Customer Focus
  • Operational Excellence
  • Communication Skills
  • Executive Presence
  • Collaboration
  • Continuous Improvement
  • Decision Making Under Pressure
Success Metrics
  • Service availability and uptime
  • Incident response times
  • Mean Time to Detect (MTTD)
  • Mean Time to Restore (MTTR)
  • Reduction in recurring incidents
  • Monitoring effectiveness
  • Automation adoption
  • Operational readiness compliance
  • Customer impact reduction
  • Service reliability improvements
Work Environment
  • 24x7 operational support organization.
  • Participation in on-call and major incident rotations.
  • Collaboration with global teams across multiple regions.
  • Hybrid cloud and enterprise production environments.
  • Fast-paced, mission-critical operational setting.

Offers of employment are conditional upon passage of screening criteria applicable to the job

EEO Statement

Integrated into our shared values is NCR Voyix's commitment to equal employment opportunity. All qualified applicants will receive consideration for employment without regard to sex, age, race, color, creed, religion, national origin, disability, sexual orientation, gender identity, veteran status, military service, genetic information, or any other characteristic or conduct protected by law.

NCR Voyix is committed to being a globally inclusive company where all people are treated fairly, recognized for their individuality, promoted based on performance and encouraged to strive to reach their full potential. We believe in understanding and respecting differences among all people. Every individual at NCR Voyix has an ongoing responsibility to respect and support a globally diverse environment.

Statement to Third Party Agencies

To ALL recruitment agencies: NCR Voyix only accepts resumes from agencies on the preferred supplier list. Please do not forward resumes to our applicant tracking system, NCR Voyix employees, or any NCR Voyix facility. NCR Voyix is not responsible for any fees or charges associated with unsolicited resumes

When applying for a job, please make sure to only open emails that you will receive during your application process that come from a @ncrvoyix.com email domain.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Site Reliability Engineer
Sr Site Reliability Engineer

NCR Corporation • India

On-site
INR 1,200,000 - 2,400,000
SW Dev Ops Security Engineer III
SW Dev Ops Security Engineer III

3M HEALTHCARE • Chennai District

On-site
INR 1,800,000 - 3,000,000
Senior SRE – Unified Observability Engineer
Senior SRE – Unified Observability Engineer

NCR Voyix • Hyderabad

On-site
INR 2,000,000 - 3,200,000
Sr Site Reliability Engineer
Sr Site Reliability Engineer

NCR Voyix • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Sr Site Reliability Engineer
Sr Site Reliability Engineer

NCR Corporation • Gurgaon

On-site
INR 900,000 - 1,300,000
Telecom Networking Engineer
Telecom Networking Engineer

NCR Corporation • Chennai District

On-site
INR 900,000 - 1,500,000
Senior Database Administrator
Senior Database Administrator

NCR Corporation • Chennai District

On-site
INR 1,800,000 - 3,800,000
Quality Engineer (II)
Quality Engineer (II)

NCR Corporation • Hyderabad

On-site
INR 900,000 - 1,400,000
Senior Site Reliability Engineer – Unified Observability
Senior Site Reliability Engineer – Unified Observability

NCR Voyix • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Telecom Networking Engineer
Telecom Networking Engineer

NCR Voyix • Chennai District

On-site
INR 600,000 - 1,200,000