Lead Site Reliability Engineer - Cloud Reliability & DR

Peraton

Northern (KY)

Hybrid

USD 112,000 - 179,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Peraton is seeking a Lead Site Reliability Engineer to join our team responsible for the reliability, resilience, and recoverability of mission-critical platforms. You will lead infrastructure-level disaster recovery drills, validate platform rebuild procedures, and oversee automated deployment processes.

You will partner with cross-functional teams to modernize cloud environments, implement IaC with Terraform and CloudFormation, manage Kubernetes clusters, and build observability using

Qualifications

  • Bachelor’s degree and 8–10 years of relevant SRE/DevOps experience; or 12 years with a high school diploma.
  • Expert level hands on knowledge of AWS services across compute, networking, storage, IAM, and serverless components.
  • Strong experience with Infrastructure as Code (Terraform, CloudFormation) and infrastructure automation principles.
  • Experience building CI/CD deployment pipelines and progressive delivery mechanisms using GitHub actions or similar tools.
  • Deep understanding of Kubernetes administration, container orchestration, and Docker based deployments.
  • Proven experience validating DR processes, performing system rebuilds, and conducting data integrity checks.

Responsibilities

  • Leading DR drills and platform rebuild validation to ensure 48-hour recovery targets are met.
  • Executing infrastructure-level drill activities to validate complete rebuild capability.
  • Verifying data integrity during drills and documenting remediation recommendations.
  • Identifying and closing readiness gaps across infra, automation, monitoring, and data recovery.
  • Designing and supporting automated IaC workflows using Terraform, CloudFormation, and CI/CD pipelines.
  • Managing and optimizing Kubernetes clusters and Docker workloads, including scaling and reliability improvements.
  • Building observability through CloudWatch, Datadog, and other monitoring tools for proactive incident response.
  • Developing automation scripts using Python/Java to reduce manual tasks.
  • Collaborating with platform engineering, security, applications, and data teams on secure, compliant operations.
  • Participating in on-call rotations and incident response to improve resilience.

Skills

AWS services
IaC (Terraform/CloudFormation)
CI/CD (GitHub Actions)
Kubernetes/Docker
Observability (CloudWatch, Datadog)
Python/Java/C#
Disaster Recovery
Incident response
Security/compliance
Documentation

Education

Bachelor’s degree
8–10 years SRE/DevOps experience

Tools

Terraform
CloudFormation
GitHub Actions
Kubernetes
Docker
CloudWatch
Datadog

Job description

Peraton is seeking a Lead Site Reliability Engineer to join our team responsible for the reliability, resilience, and recoverability of mission-critical platforms. You will lead infrastructure-level disaster recovery drills, validate platform rebuild procedures, and oversee automated deployment processes.

You will partner with cross-functional teams to modernize cloud environments, implement IaC with Terraform and CloudFormation, manage Kubernetes clusters, and build observability using

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer: DR, Cloud & Automation Leader
Site Reliability Engineer: DR, Cloud & Automation Leader

Peraton • Reston (VA)

On-site
USD 112,000 - 179,000
Senior Site Reliability Engineer — AWS, Kubernetes & DR
Senior Site Reliability Engineer — AWS, Kubernetes & DR

Peraton • United States

Remote
USD 130,000 - 180,000
Senior Site Reliability Engineer: DR, Cloud & Automation Lead
Senior Site Reliability Engineer: DR, Cloud & Automation Lead

Peraton • Herndon (VA)

On-site
USD 112,000 - 179,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Peraton • United States

Remote
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Lead Site Reliability Engineer - Architect & Own Production
Lead Site Reliability Engineer - Architect & Own Production

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000
Senior SRE Lead: Cloud Reliability & Automation
Senior SRE Lead: Cloud Reliability & Automation

Oracle • Honolulu (HI)

On-site
USD 96,000 - 264,000
Medical, dental, and vision insurance
Paid time off
401(k) with company match
+4
Remote Lead SRE - Multi-Cloud & Automation
Remote Lead SRE - Multi-Cloud & Automation

Koitecc Solutions • San Antonio (TX)

Hybrid
USD 116,000 - 210,000
Site Reliability Engineer - Build Resilient Cloud Systems
Site Reliability Engineer - Build Resilient Cloud Systems

Avalore • Virginia (MN)

On-site
USD 107,000 - 220,000
Health care plan
401k/IRA with matching
Life Insurance
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ConsultNet Technology Services and Solutions • El Segundo (CA)

On-site
USD 140,000 - 180,000