Production Reliability Engineer (SRE & Automation)

NRnP Technology

Northern (KY)

Hybrid

USD 90,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NRnP Technology is seeking an experienced Production Support Engineer to join our 24/7 operations team in the United States. You will act as the primary contact for incidents, perform triage, root cause analysis, and code-level fixes, and collaborate across engineering, QA, and infra to minimize business impact.

You will design automation to reduce manual tasks, build observability dashboards, participate in post-incident reviews, and help with capacity planning.

Qualifications

  • 3+ years of experience in application support, production support, or a similar engineering role.
  • Strong understanding of software engineering principles and ability to read and debug code.
  • Experience with monitoring and observability tools.
  • Familiarity with cloud platforms and containerized environments.
  • Experience building automation scripts using Python or similar languages.
  • Familiarity with CI/CD pipelines and DevOps practices.
  • Strong problem-solving and analytical skills.
  • Excellent verbal and written communication skills.

Responsibilities

  • Act as the primary point of contact for production issues, alerts, and user-reported incidents.
  • Log, categorize, and prioritize tickets in accordance with defined SLAs.
  • Ensure timely resolution and minimize business impact.
  • Perform root cause analysis for production issues across application and infrastructure layers.
  • Collaborate with cross-functional teams to resolve complex technical problems.
  • Use logs, metrics, and traces to diagnose issues efficiently.
  • Review and debug application code to identify and resolve defects.
  • Implement code-level fixes and coordinate with development teams for larger changes.
  • Participate in code reviews to ensure quality and maintainability.
  • Design and build automation scripts and tools to reduce manual operational tasks.
  • Leverage AI-assisted tools to streamline troubleshooting and incident response.
  • Continuously identify opportunities to improve operational workflows.
  • Build and maintain dashboards, alerts, and monitoring frameworks.
  • Enhance system observability to proactively detect and prevent issues.
  • Ensure monitoring coverage aligns with business-critical services.
  • Apply SRE principles to improve system reliability, availability, and performance.
  • Participate in post-incident reviews and drive corrective actions.
  • Contribute to capacity planning and performance tuning efforts.
  • Escalate unresolved issues to appropriate teams while maintaining ownership of resolution.
  • Work closely with development, QA, and infrastructure teams to prevent recurring issues.
  • Foster strong cross-team relationships to support rapid issue resolution.
  • Create and maintain runbooks, troubleshooting guides, and knowledge base articles.
  • Document incident details, root causes, and resolutions for future reference.
  • Ensure documentation remains current and accessible to the team.
  • Provide clear and timely updates to stakeholders during incidents.
  • Communicate technical issues effectively to both technical and non-technical audiences.
  • Collaborate with global teams across different time zones.

Skills

Application support
Observability
Coding skills
Problem solving
Communication

Education

Bachelor’s degree in Computer Science, Engineering, or related field

Tools

Docker
Kubernetes
Python

Job description

NRnP Technology is seeking an experienced Production Support Engineer to join our 24/7 operations team in the United States. You will act as the primary contact for incidents, perform triage, root cause analysis, and code-level fixes, and collaborate across engineering, QA, and infra to minimize business impact.

You will design automation to reduce manual tasks, build observability dashboards, participate in post-incident reviews, and help with capacity planning.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production SRE: Incident Response & Automation
Production SRE: Incident Response & Automation

Mobility Global • Michigan

On-site
USD 85,000 - 125,000
Senior SRE: Production Reliability & Automation
Senior SRE: Production Reliability & Automation

PNC Financial Services Group, Inc. • Denver (CO)

On-site
USD 86,000 - 158,000
Medical and prescription drug coverage
Dental and vision
401(k) with company match
+4
Senior SRE: Stabilize & Automate Production
Senior SRE: Stabilize & Automate Production

PNC • Cleveland (OH)

On-site
USD 86,000 - 158,000
Medical coverage
Dental & vision options
401(k) with company match
+2
SRE Engineering Manager — 24x7 Production Reliability
SRE Engineering Manager — 24x7 Production Reliability

PNC • Pittsburgh

On-site
USD 115,000 - 150,000
Medical and prescription drug coverage
401(k) with PNC match
Paid time off and holiday benefits
+1
Remote SRE Manager: Lead Reliability & Automation
Remote SRE Manager: Lead Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
SRE Engineering Manager — On-Call Production Leader
SRE Engineering Manager — On-Call Production Leader

PNC • Phoenix (AZ)

On-site
USD 100,000 - 204,000
Health insurance
401(k) with match
Paid time off
SRE Production Support Engineer | 24x7 Reliability & DR
SRE Production Support Engineer | 24x7 Reliability & DR

Micasa Global • Vienna (VA)

Hybrid
USD 100,000 - 130,000
SRE: Production Reliability Engineer (Contract)
SRE: Production Reliability Engineer (Contract)

Tech Mirrors • Pittsburgh

On-site
USD 110,000 - 150,000
Production Stability & Incident Engineer (SRE)
Production Stability & Incident Engineer (SRE)

Standard Chartered • United States

Hybrid
USD 120,000 - 160,000
Flexible working options
Competitive pay and benefits
Annual/ parental leave and sabbatical
+3
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000