Site Reliability Engineer (SRE) – Evening Shift

Peraton

Northern (KY)

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Peraton is seeking a Site Reliability Engineer (SRE) to join a team responsible for the operational reliability of production systems running in AWS Commercial and AWS GovCloud environments. The ideal candidate has working knowledge of deploying and managing system components in Azure and GCP.

The primary production workload runs on Red Hat OpenShift Service on AWS (ROSA) The SRE partners closely with platform engineers, the security team, and application developers to ensure the infrastructure

Qualifications

  • Bachelor's Degree or equivalent experience with 8+ years in related fields
  • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering
  • Experience operating in AWS Commercial and GovCloud, including OpenShift ROSA or similar Kubernetes platforms
  • Proficient with infrastructure-as-code tools Terraform and Ansible/Ansible Tower
  • Experience with CI/CD platforms GitLab and Jenkins and deployment automation
  • Proficient in Linux and Windows Server administration
  • Experience with enterprise observability tools like Dynatrace, Datadog, Splunk, OpenTelemetry
  • Demonstrated ownership of SLI/SLO and alerting programs
  • Scripting in Python, Bash, PowerShell, or Go
  • Experience in federal or regulated environments (FISMA, FedRAMP, NIST 800-53)

Responsibilities

  • Operate and maintain production infrastructure services ensuring availability, reliability, performance, security, and health
  • Monitor services using SLIs/SLOs, dashboards, alerts; improve issue detection and resolution
  • Collaborate with application teams to define observability requirements and implement metrics, logs, traces, dashboards, and alerts
  • Manage production incidents including on-call response, troubleshooting, root-cause analysis, and post-incident actions
  • Release applications and infrastructure through pipelines with staging/production promotion, validation, and rollback
  • Manage lifecycle of deployed infrastructure including upgrades, patches, and maintenance
  • Assess and improve service resilience through capacity planning, testing, disaster recovery, and backups
  • Identify reliability risks and tech debt using metrics and incident trends to drive improvements
  • Automate operational activities via as-code approach for consistency and efficiency
  • Collaborate with platform engineers and developers to improve reliability and operability

Skills

Site Reliability Engineering
DevOps
Production systems engineering
AWS GovCloud/AWS Commercial
OpenShift ROSA/Kubernetes
Terraform/Ansible
CI/CD (GitLab/Jenkins)
Linux/Windows Server
Observability tools (Dynatrace/Datadog
SLI/SLO and alerting
Python/Bash/PowerShell/Go
Security/compliance (FISMA/FedRAMP/NIS

Education

Bachelor's Degree or equivalent experience

Tools

OpenShift ROSA
Terraform
Ansible/Ansible Tower
GitLab
Jenkins

Job description

Required Qualifications:
  • Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance.
  • Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience.
  • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
  • Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms
  • Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.
  • Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation.
  • Proficient in Linux and Windows Server administration
  • Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry.
  • Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction.
  • Scripting/automation proficiency in Python, Bash, PowerShell, or Go.
  • Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).
Preferred Qualifications:
  • AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification
  • Red Hat Certified Specialist in ROSA, Red Hat Certified System Administrator in OpenShift
  • Azure Administrator Associate, GCP Associate Cloud Engineer certification
  • Dynatrace Associate, Datadog Log Management Fundamentals certification
  • GitLab CI/CD Associate certification, Certified Jenkins Engineer (CJE)
  • Terraform Associate certification

Peraton is seeking a Site Reliability Engineer (SRE) to join a team respobsible for the operational reliability of production systems running in AWS Commercial and AWS GovCloud environments. The ideal candidate has working knowledge of deploying and managing system components in Azure and GCP. The primary production workload runs on Red Hat OpenShift Service on AWS (ROSA)

The SRE partners closely with platform engineers, the security team, and application developers to ensure the infrastructure services are reliable, available, and deployed in a way that meets both developer and security requirements.

Work Location: Remote

Shift Schedule: This is an EveningShift position with working hours from 3pm - 11pm Eastern Standard Time (EST)

What you will do:
  • Operate and maintain production infrastructure services and applications, ensuring availability and reliability, performance, security, and operational health.
  • Monitor services and applications using defined SLIs, SLOs, dashboards, alerts, and other observability tools; continuously improve the detection, diagnosis, and resolution of operational issues.
  • Partner with application teams to define application observability requirements and implement appropriate metrics, logs, traces, dashboards, and alerts into the organization's observability tooling.
  • Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions.
  • Execute application and infrastructure releases through established deployment pipelines, including promotion through staging and production, validation, rollback, and release-related troubleshooting.
  • Manage the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes.
  • Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.
  • Identify and address reliability risks and operational technical debt by using reliability metrics, incident trends, capacity data, and service health indicators to prioritize improvements.
  • Automate operational activities using an everything-as-code approach to improve consistency, repeatability, testing, deployment, recovery, and operational efficiency.
  • Collaborate with platform engineering and application teams to identify operational requirements, provide feedback on reusable infrastructure building blocks, and continuously improve the reliability and operability of the environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
External Job Posting Title Site Reliability Engineer (SRE) – Night Shift
External Job Posting Title Site Reliability Engineer (SRE) – Night Shift

Peraton • Northern (KY)

Hybrid
USD 104,000 - 166,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Site Reliability Engineer (Secret Clearance)
Site Reliability Engineer (Secret Clearance)

ROI Services LLC • Huntsville (AL)

On-site
USD 110,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Chicago (IL)

On-site
USD 120,000 - 180,000
Evening SRE – Remote Cloud & OpenShift Reliability Engineer
Evening SRE – Remote Cloud & OpenShift Reliability Engineer

Peraton • Reston (VA)

On-site
USD 104,000 - 166,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Remote SRE - AWS/OpenShift Reliability Engineer (Evening)
Remote SRE - AWS/OpenShift Reliability Engineer (Evening)

Peraton • Northern (KY)

Hybrid
USD 120,000 - 180,000