Site Reliability Engineer II (NGPOS Operations Support)

hudsonmanpower

Cincinnati (OH)

On-site

USD 110,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

HudsonManpower is seeking a hands-on Site Reliability Engineer II to support a next-generation POS platform in a production environment. The role emphasizes production reliability, incident leadership, and operational excellence in a fast-paced setting.

The ideal candidate will lead major incident responses, drive RCA, and collaborate with Software, Platform, Infrastructure, and Business Operations teams to improve platform reliability.

Qualifications

  • 3+ years of experience in Site Reliability Engineering or a related field.
  • Strong incident management and RCA skills.
  • Excellent communication under pressure and ownership.
  • Experience coordinating multiple engineering teams and driving operational excellence.

Responsibilities

  • Lead major incident response during production outages.
  • Serve as Incident Commander during P1/P2 incidents.
  • Coordinate technical bridge calls and communicate outage status to teams and leadership.
  • Lead Root Cause Analysis (RCA) activities and track corrective actions.
  • Improve production reliability through better monitoring, observability, and runbooks.

Skills

Incident management
Incident Commander
Cross-functional coordination
Observability
Troubleshooting
Leadership
Communication under pressure
On-call rotations

Tools

Dynatrace
Azure Monitor
Log Analytics
Kubernetes
Docker
Jira

Job description

Position Overview

We are seeking a hands‑on Site Reliability Engineer II to support a next-generation Point of Sale (NGPOS) platform in a highly visible production environment. Unlike traditional SRE roles focused primarily on automation or platform engineering, this position emphasizes production reliability, incident leadership, operational excellence, and engineering support.

The ideal candidate will lead major incident response efforts, drive Root Cause Analysis (RCA), improve system observability, and collaborate closely with Software Engineering, Platform Engineering, Infrastructure, and Business Operations teams to enhance overall platform reliability.

This role is ideal for someone who enjoys solving complex production issues under pressure while contributing to long‑term engineering improvements.

Location and Employment Type

Location: Blue Ash, OH (Cincinnati) – Onsite (5 Days/Week)

Employment Type: W2 – Contract‑to‑Hire

Duration: Full‑Time

Work Authorization: Permanent Residents only. Must be able to convert to full‑time without sponsorship.

Experience Required: 3+ Years

Nice to Have
  • Enterprise Point of Sale (POS) systems
  • Retail technology experience
  • Automation scripting
  • Monitoring optimization
  • Runbook creation
  • Store technology deployments
Key Responsibilities
  • Lead major incident response during production outages
  • Serve as Incident Commander during P1/P2 incidents
  • Coordinate technical bridge calls
  • Communicate outage status to engineering teams and business leadership
  • Lead Root Cause Analysis (RCA) activities
  • Track corrective actions through completion
  • Improve production reliability and system stability
  • Enhance monitoring and observability
  • Reduce alert fatigue
  • Partner with Software Engineering and Platform Engineering teams
  • Support retail store deployments
  • Develop operational documentation, runbooks, and playbooks
  • Participate in after‑hours support rotations and maintenance windows
  • Improve service health using SLIs and SLOs
Technical Environment

Monitoring & Observability

  • Dynatrace
  • Azure Monitor
  • Log Analytics
  • Metrics
  • Dashboards

Cloud

  • Microsoft Azure
  • Google Cloud Platform (GCP)

Containers

  • Kubernetes
  • Docker

Operating Systems

  • Linux

Scripting Languages

  • Bash
  • Python

Agile Tools

  • Jira

Enterprise Environment

  • Retail systems
  • Point of Sale (POS)
  • Production Support
  • Hybrid Infrastructure
Ideal Candidate Profile

The ideal candidate will demonstrate:

  • Strong leadership during production incidents
  • Excellent troubleshooting and analytical skills
  • Effective communication under pressure
  • Ownership and accountability
  • Experience coordinating multiple engineering teams
  • Strong operational discipline
  • Continuous improvement mindset
  • Passion for reliability engineering
  • Excellent documentation skills
Required Skills

Must Have

  • Major Incident Management experience
  • Experience serving as Incident Commander
  • Leading production bridge calls
  • Coordinating cross- functional technical teams
  • Executive communication during P1/P2 outages
  • Root Cause Analysis (RCA)
    • Five Whys
    • Fishbone Analysis
    • Timeline reconstruction
    • Corrective action tracking
  • Production troubleshooting across:
    • Cloud environments
    • On-premise infrastructure
    • Retail/POS systems
  • Observability and Monitoring
    • Dynatrace
    • Azure Monitor
    • Log analysis
    • Dashboards
    • Metrics
  • Linux administration
  • Bash and/or Python scripting
  • Kubernetes
  • Docker
  • Microsoft Azure and/or Google Cloud Platform (GCP)
  • Agile methodology
  • Jira
  • Strong communication and collaboration skills
  • Ability to work onsite five days per week
  • Willingness to travel and participate in on-call rotations
Travel Requirements
  • Local travel initially throughout Cincinnati and Louisville
  • Travel will expand as additional retail locations are deployed
  • Approximately every four weeks as new sites go live
  • Shared on-call and travel rotation with a growing six‑person team
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Support Engineer
Support Engineer

Flexton Inc. • Blue Ash (OH)

On-site
USD 95,000 - 130,000
Technical Support Engineer
Technical Support Engineer

Ascendum Solutions • Cincinnati (OH)

On-site
USD 85,000 - 110,000
I Site Reliability Engineer Incident IQ Alpharetta, Georgia, US
I Site Reliability Engineer Incident IQ Alpharetta, Georgia, US

Artha Nexgen • Alpharetta (GA), Northern (KY)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

Hybrid
USD 120,000 - 150,000
Site-Reliability Engineer, Application Operations
Site-Reliability Engineer, Application Operations

Scorpion Therapeutics • Irving (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Koitecc Solutions • Chandler (AZ), Northern (KY)

Hybrid
USD 140,000 - 200,000
Retail POS Application Support
Retail POS Application Support

City of Lincoln • Westlake (OH)

On-site
USD 41,000 - 45,000
Onsite SRE II: Production Reliability & Incident Leader
Onsite SRE II: Production Reliability & Incident Leader

hudsonmanpower • Cincinnati (OH)

On-site
USD 110,000 - 140,000
Platform Engineer & Production Support
Platform Engineer & Production Support

Strategic Staffing Solutions • Charlotte (NC)

Hybrid
USD 100,000 - 150,000