Site Reliability Engineer (SRE) - Night Shift

Peraton

United States

On-site

USD 104,000 - 166,000

Full time

3 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical benefits
Dental benefits
Vision benefits
401(k) plan
Paid time off
Parental leave

Job summary

Peraton is seeking a Site Reliability Engineer (SRE) to join a team responsible for the operational reliability of production systems running in AWS Commercial and AWS GovCloud environments. The role requires OpenShift (ROSA) or Kubernetes experience and a shift that aligns with 11pm–7am EST hours.

The position emphasizes on-call incident management, automation, and observability, with a focus on reliability, capacity planning, and secure operations in federal/regulatory contexts.

Qualifications

  • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
  • Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms.
  • Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.
  • Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation.
  • Proficient in Linux and Windows Server administration.
  • Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry.
  • Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction.

Responsibilities

  • Operate and maintain production infrastructure services and applications, ensuring availability and reliability, performance, security, and operational health.
  • Monitor services and applications using SLIs/SLOs, dashboards, alerts, and observability tools; continuously improve detection, diagnosis, and resolution of issues.
  • Collaborate with platform engineering and application teams to define observability requirements and implement metrics, logs, traces, and alerts.
  • Manage production incidents and on-call response, troubleshooting, restoration, root-cause analysis, and post-incident actions.
  • Execute releases through deployment pipelines, including staging/production promotion, validation, rollback, and troubleshooting.
  • Automate operational activities using an everything-as-code approach to improve consistency and efficiency.
  • Support capacity planning, disaster recovery, backup, failover, and service resilience initiatives.

Skills

SRE/DevOps experience
Automation & scripting
On-call incident handling
OpenShift/Kubernetes familiarity
Cloud service experience

Education

Bachelor's Degree + 8 years experience
High School diploma + 12 years experience

Tools

Terraform
Ansible
GitLab
Jenkins
Dynatrace
Datadog
Splunk
OpenTelemetry
AWS/GovCloud

Job description

About Peraton

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world's leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees solve the most daunting challenges that our customers face. Visit peraton.com to learn how we're keeping people around the world safe and secure.

About Peraton

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world's leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees solve the most daunting challenges that our customers face. Visit peraton.com to learn how we're keeping people around the world safe and secure.

About The Role

Peraton is seeking a Site Reliability Engineer (SRE) to join a team respobsible for the operational reliability of production systems running in AWS Commercial and AWS GovCloud environments. The ideal candidate has working knowledge of deploying and managing system components in Azure and GCP. The primary production workload runs on Red Hat OpenShift Service on AWS (ROSA)

Work Location: Remote

Shift Schedule: This is a Night Shift position with working hours from 11pm - 7am Eastern Standard Time (EST)

What you will do:
  • Operate and maintain production infrastructure services and applications, ensuring availability and reliability, performance, security, and operational health.
  • Monitor services and applications using defined SLIs, SLOs, dashboards, alerts, and other observability tools; continuously improve the detection, diagnosis, and resolution of operational issues.
  • Partner with application teams to define application observability requirements and implement appropriate metrics, logs, traces, dashboards, and alerts into the organization's observability tooling.
  • Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions.
  • Execute application and infrastructure releases through established deployment pipelines, including promotion through staging and production, validation, rollback, and release-related troubleshooting.
  • Manage the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes.
  • Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.
  • Identify and address reliability risks and operational technical debt by using reliability metrics, incident trends, capacity data, and service health indicators to prioritize improvements.
  • Automate operational activities using an everything-as-code approach to improve consistency, repeatability, testing, deployment, recovery, and operational efficiency.
  • Collaborate with platform engineering and application teams to identify operational requirements, provide feedback on reusable infrastructure building blocks, and continuously improve the reliability and operability of the environment.
Qualifications
Basic Qualifications
  • Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance.
  • Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience.
  • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
  • Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms
  • Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.
  • Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation.
  • Proficient in Linux and Windows Server administration
  • Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry.
  • Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction.
  • Scripting/automation proficiency in Python, Bash, PowerShell, or Go.
  • Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).
Preferred Qualifications
  • AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification
  • Red Hat Certified Specialist in ROSA, Red Hat Certified System Administrator in OpenShift
  • Azure Administrator Associate, GCP Associate Cloud Engineer certification
  • Dynatrace Associate, Datadog Log Management Fundamentals certification
  • GitLab CI/CD Associate certification, Certified Jenkins Engineer (CJE)
  • Terraform Associate certification
Details

Target Salary Range: $104,000 - $166,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual's experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

Benefits Statement: Peraton offers eligible employees a variety of benefits including medical, dental, vision, life, health savings account, short/long term disability, EAP, parental leave, 401(k), paid time off (PTO) for vacation, and company paid holidays. A full listing of available benefits can be viewed at https://www.careers.peraton.com/benefits.

Application Statements: The application period for the job is estimated to be 30 days from the job posting date. However, this timeline may be shortened or extended depending on business needs and the availability of qualified candidates. By applying to this job, you are expressing interest in the role and the Company. During the review of your application, you may be required to participate in an on-camera interview, as well as participate in a process to verify your identity. Use of artificial intelligence (AI) tools of any kind during Peraton interviews is strictly prohibited unless the candidate has obtained prior written authorization. All interview responses must be the candidate's own.

EEO:Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

External Job Posting Title Site Reliability Engineer (SRE) – Evening Shift
External Job Posting Title Site Reliability Engineer (SRE) – Evening Shift

Peraton • Northern (KY)

Hybrid
USD 104,000 - 166,000
External Job Posting Title Site Reliability Engineer (SRE) – Night Shift
External Job Posting Title Site Reliability Engineer (SRE) – Night Shift

Peraton • Northern (KY)

Hybrid
USD 104,000 - 166,000
External Job Posting Title Site Reliability Engineer
External Job Posting Title Site Reliability Engineer

Peraton • Washington

On-site
USD 112,000 - 179,000
Site Reliability Engineer
Site Reliability Engineer

Peraton • United States

On-site
USD 112,000 - 179,000
Competitive benefits
Site Reliability Engineer, Senior Advisor
Site Reliability Engineer, Senior Advisor

Peraton • United States

On-site
USD 190,000 - 304,000
Generous PTO
Subsidized medical, dental, and vision
Competitive bonus plan
External Job Posting Title Senior Platform Engineer/Application Enablement Lead
External Job Posting Title Senior Platform Engineer/Application Enablement Lead

Peraton • Northern (KY)

Hybrid
USD 112,000 - 179,000
Technical Enterprise Incident Manager
Technical Enterprise Incident Manager

Peraton • United States

On-site
USD 86,000 - 138,000
Medical and dental insurance
401(k) plan
Paid time off (PTO)
External Job Posting Title Lead Cloud Architect
External Job Posting Title Lead Cloud Architect

Peraton • Northern (KY)

Hybrid
USD 112,000 - 179,000
Site Reliability - Java Springboot Applications
Site Reliability - Java Springboot Applications

Peraton • Herndon (VA)

On-site
USD 112,000 - 179,000
25 days PTO
Bonuses eligible
Dependent benefits
Systems Engineering, TS/SCI w/Poly
Systems Engineering, TS/SCI w/Poly

Peraton • Maryland

On-site
USD 146,000 - 234,000
Health benefits
PTO 25 days per year
Bonus plan eligibility