Remote Night-Shift SRE: AWS/GovCloud Reliability Engineer

Peraton

Northern (KY)

Hybrid

USD 104,000 - 166,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Peraton is seeking a Site Reliability Engineer (SRE) to join a team responsible for the operational reliability of production systems running in AWS Commercial and AWS GovCloud environments, including OpenShift ROSA. The role partners with platform engineers, security, and application teams to ensure reliable infrastructure aligned with developer and security requirements.

The position supports a night shift from 11pm–7am EST and emphasizes proactive monitoring, automation, and incident response

Qualifications

  • U.S. Citizenship with the ability to obtain and maintain the required Public Trust clearance.
  • Bachelor's Degree with 8+ years of experience, or HS diploma/equivalent with 12+ years of experience.
  • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
  • Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms.
  • Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.
  • Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation.
  • Proficient in Linux and Windows Server administration.
  • Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry.
  • Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction.
  • Scripting/automation proficiency in Python, Bash, PowerShell, or Go.
  • Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).

Responsibilities

  • Operate and maintain production infrastructure services and applications, ensuring availability and reliability, performance, security, and operational health.
  • Monitor services and applications using defined SLIs, SLOs, dashboards, alerts, and other observability tools; continuously improve the detection, diagnosis, and resolution of operational issues.
  • Partner with application teams to define application observability requirements and implement appropriate metrics, logs, traces, dashboards, and alerts into the organization's observability tooling.
  • Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions.
  • Execute application and infrastructure releases through established deployment pipelines, including promotion through staging and production, validation, rollback, and release-related troubleshooting.
  • Manage the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes.
  • Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.
  • Identify and address reliability risks and operational technical debt by using reliability metrics, incident trends, capacity data, and service health indicators to prioritize improvements.
  • Automate operational activities using an everything-as-code approach to improve consistency, repeatability, testing, deployment, recovery, and operational efficiency.
  • Collaborate with platform engineering and application teams to identify operational requirements, provide feedback on reusable infrastructure building blocks, and continuously improve the reliability and operability of the environment.

Skills

Site Reliability Engineering
DevOps
Scripting
Linux & Windows
OpenShift ROSA / Kubernetes
Terraform
Ansible
CI/CD
Observability

Education

Bachelor's Degree
High School Diploma / equivalent

Tools

Terraform
Ansible/Ansible Tower
GitLab
Jenkins
Dynatrace
Datadog
Splunk
OpenTelemetry

Job description

Peraton is seeking a Site Reliability Engineer (SRE) to join a team responsible for the operational reliability of production systems running in AWS Commercial and AWS GovCloud environments, including OpenShift ROSA. The role partners with platform engineers, security, and application teams to ensure reliable infrastructure aligned with developer and security requirements.

The position supports a night shift from 11pm–7am EST and emphasizes proactive monitoring, automation, and incident response

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Night-Shift SRE — Cloud Infra Reliability
Remote Night-Shift SRE — Cloud Infra Reliability

Peraton • Reston (VA)

On-site
USD 104,000 - 166,000
Evening SRE – Remote Cloud & OpenShift Reliability Engineer
Evening SRE – Remote Cloud & OpenShift Reliability Engineer

Peraton • Reston (VA)

On-site
USD 104,000 - 166,000
Evening SRE - Remote Cloud Reliability Engineer (AWS/ROSA)
Evening SRE - Remote Cloud Reliability Engineer (AWS/ROSA)

Peraton • Herndon (VA)

On-site
USD 104,000 - 166,000
Night-Shift SRE (Remote) – AWS/ROSA & Observability
Night-Shift SRE (Remote) – AWS/ROSA & Observability

Peraton • Herndon (VA)

On-site
USD 104,000 - 166,000
Remote Evening SRE: ROSA & AWS Observability Expert
Remote Evening SRE: ROSA & AWS Observability Expert

Peraton • Northern (KY)

Hybrid
USD 104,000 - 166,000
External Job Posting Title Site Reliability Engineer (SRE) – Night Shift
External Job Posting Title Site Reliability Engineer (SRE) – Night Shift

Peraton • Northern (KY)

Hybrid
USD 104,000 - 166,000
External Job Posting Title Site Reliability Engineer (SRE) – Evening Shift
External Job Posting Title Site Reliability Engineer (SRE) – Evening Shift

Peraton • Northern (KY)

Hybrid
USD 104,000 - 166,000
Remote Senior SRE — Cloud Reliability & Automation
Remote Senior SRE — Cloud Reliability & Automation

EverCommerce • Denver (CO)

Hybrid
USD 110,000 - 130,000
Flexible work environment
Health and wellness benefits
401(k) with company match
+2
Remote Lead Cloud Architect – Reliability & Automation
Remote Lead Cloud Architect – Reliability & Automation

Peraton • Herndon (VA)

On-site
USD 112,000 - 179,000
Senior SRE - Cloud Infra, Automation & 24/7 Ops
Senior SRE - Cloud Infra, Automation & 24/7 Ops

Oracle • Reston (VA)

On-site
USD 85,000 - 210,000
Medical, dental, and vision insurance
Paid time off
401(k) Savings Plan