Lead Site Reliability Engineer

Peraton

Herndon (VA)

On-site

USD 112,000 - 179,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Peraton in Herndon, VA seeks a Lead Site Reliability Engineer to ensure reliability and recoverability of mission-critical platforms. You will lead DR drills, validate platform rebuilds, and drive automation across IaC and CI/CD pipelines.

You will manage Kubernetes clusters, container workloads, and observability via CloudWatch and Datadog, while collaborating with cross-functional teams to improve resilience and on-call readiness.

Qualifications

  • 8–10 years of relevant SRE/DevOps/infrastructure experience; or 12+ years with a high school diploma.
  • Expert level hands-on knowledge of AWS services across compute, networking, storage, IAM, and serverless components.
  • Strong experience with Infrastructure as Code (Terraform, CloudFormation) and automation principles.
  • Experience building CI/CD pipelines and progressive delivery with GitHub Actions or similar tools.
  • Deep understanding of Kubernetes administration, container orchestration, and Docker deployments.
  • Proven experience validating DR processes, rebuilding systems, and data integrity checks.
  • Experience building monitoring tools using CloudWatch, Datadog, or similar observability tools.
  • Proficiency in Python, Java, C#, or Go for automation.
  • Experience debugging complex failure modes and cascading issues.
  • Strong analytical and documentation skills; able to communicate findings clearly.
  • Ability to work in a fast-paced environment on mission-critical systems; eligible for Public Trust clearance.
  • US Citizen or Green Card Holder.

Responsibilities

  • Lead lifecycle platform portability and DR drill execution, including validating rebuild procedures and DR playbooks.
  • Execute infrastructure-level drills to ensure recovery within the 48-hour target.
  • Verify data integrity during drills and document remediation recommendations.
  • Identify exit readiness gaps across infrastructure, automation, monitoring, and data recovery; drive corrective actions.
  • Design and implement automated IaC workflows with Terraform and CloudFormation; standardize CI/CD pipelines.
  • Manage and optimize Kubernetes clusters and container workloads (Docker) including provisioning and scaling.
  • Build and maintain observability solutions using CloudWatch, Datadog, and related tools.
  • Develop automation scripts in Python or Java to reduce manual work and improve repeatability.
  • Collaborate with platform engineering, security, applications, and data teams for compliant operations.
  • Participate in on-call rotations, root cause analyses, and incident response to improve resilience.

Skills

AWS
Terraform
CloudFormation
CI/CD
Kubernetes
Docker
Python
Java
Public Trust clearance
US Citizenship

Education

Bachelor's degree

Tools

GitHub Actions
Jenkins

Job description

Job Locations

US

Responsibilities

Peraton is seeking a Lead Site Reliability Engineer to join our team of qualified, diverse individuals. The ideal candidate will play a critical role in ensuring the reliability, resilience, and recoverability of mission essential platforms by leading infrastructure level disaster recovery drills, platform rebuild validation, and automated deployment processes. This engineer will partner closely with cross functional teams to maintain and enhance complex cloud-based environments, integrate modern automation solutions, and support largescale modernization and continuity efforts across high visibility programs.

  • Supporting full lifecycle platform portability and disaster recovery (DR) drill execution, including validation of platform rebuild procedures and DR playbooks.
  • Executing infrastructure level drill activities to ensure the platform can be fully rebuilt within the 48hour recovery target.
  • Verifying end to end data completeness, integrity, and accuracy during drill exercises, documenting results and remediation recommendations.
  • Identifying exit readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes, and driving corrective actions with engineering teams.
  • Designing, implementing, and supporting automated IaaC workflows utilizing Terraform, AWS CloudFormation, and standardized CI/CD pipelines.
  • Managing and optimizing Kubernetes clusters and containerized workloads (Docker), including cluster provisioning, scaling, and workload reliability improvements.
  • Building and maintaining observability solutions using CloudWatch, Datadog, and other monitoring/alerting tools to ensure service reliability and proactive incident response.
  • Developing automation, tooling, and scripts using Python or Java to reduce manual processes and enhance operational repeatability.
  • Collaborating with platform engineering, security, applications, and data teams to ensure consistent, secure, and compliant platform operations.
  • Participating in oncall rotations, root cause analyses, and incident response activities to improve system resilience and operational excellence.
Qualifications
Required Qualifications
  • Bachelor's degree and 8-10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; or 12 years of experience with a high school diploma.
  • Expert level hands on knowledge of AWS services across compute, networking, storage, IAM, and serverless components.
  • Strong experience with Infrastructure as Code (Terraform, CloudFormation) and infrastructure automation principles.
  • Experience building CI/CD deployment pipelines and progressive delivery mechanisms using GitHub actions or similar tools
  • Deep understanding of Kubernetes administration, container orchestration, and Docker based deployments.
  • Proven experience validating DR processes, performing system rebuilds, and conducting data integrity checks.
  • Experience building monitoring tools like dashboards, metrics, logs, and alerting systems using CloudWatch, Datadog, or similar observability tools.
  • Proficiency with programming/scripting languages such as Python, Java, or C# or Go.
  • Experience debugging complex failure modes, including cascading failures, network partitions, backpressure, and eventual consistency issues.
  • Strong analytical and documentation skills with the ability to clearly communicate technical findings to cross functional teams.
  • Ability to work in a fast-paced environment supporting high visibility, mission critical systems.
  • Ability to obtain a Public Trust clearance.
  • US Citizen or Green Card Holder.
Preferred Qualifications
  • AWS DevSecOps Engineer certification (preferred).
  • Additional AWS certifications (Solutions Architect, SysOps, Developer) and/or Kubernetes certifications (CKA, CKAD).
  • Familiarity with Zero Trust security models and cloud security best practices.
  • Experience with GitLab, Jenkins, or similar CI/CD platforms.
  • Experience with highly regulated environments (healthcare, finance, DHS, DoD, CMS, etc.).
  • Experience supporting federal, defense, or largescale enterprise programs involving legacy-to-cloud modernization.
  • Prior involvement in largescale DR drills, continuity of operations (COOP), or portability/executable readiness assessments.
Peraton Overview

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world's leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can't be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we're keeping people around the world safe and secure.

Target Salary Range

$112,000 - $179,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual's experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

EEO

EEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Peraton • Reston (VA)

On-site
USD 112,000 - 179,000
External Job Posting Title Lead Site Reliability Engineer
External Job Posting Title Lead Site Reliability Engineer

Peraton • Northern (KY)

Hybrid
USD 112,000 - 179,000
External Job Posting Title Site Reliability Engineer
External Job Posting Title Site Reliability Engineer

Peraton • Washington

On-site
USD 112,000 - 179,000
External Job Posting Title Cloud Platform Technical Lead
External Job Posting Title Cloud Platform Technical Lead

Peraton • Northern (KY)

On-site
USD 112,000 - 179,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Peraton • United States

Remote
USD 130,000 - 180,000
Linux Systems Administrator (RHEL)
Linux Systems Administrator (RHEL)

Peraton • Chantilly (VA)

On-site
USD 135,000 - 216,000
PTO 25 days
Enhanced benefits
Bonus plan eligible
Software Engineer - Cloud, Lead Associate - TS/SCI w/poly
Software Engineer - Cloud, Lead Associate - TS/SCI w/poly

Peraton • Laurel (MD)

On-site
USD 176,000 - 282,000
Heavily subsidized medical, dental, &/
Vision coverage for employees & depend
25 days PTO annually
+2
Cloud Platform Technical Lead
Cloud Platform Technical Lead

Peraton • Reston (VA)

Remote
USD 112,000 - 179,000
Site Reliability - Java Springboot Applications
Site Reliability - Java Springboot Applications

Peraton • Herndon (VA)

On-site
USD 112,000 - 179,000
25 days PTO
Bonuses eligible
Dependent benefits
Test Engineering, Associate
Test Engineering, Associate

Peraton • United States

Remote
USD 66,000 - 106,000