Senior Site Reliability Engineer

GDH

Hampton (VA)

On-site

USD 94,000 - 99,000

Full time

12 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

On-site work

Job summary

GDH is seeking a Senior Site Reliability Engineer to maintain the reliability, resilience, and recoverability of mission-critical cloud platforms. You will lead DR drills, validate platform rebuilds, and automate deployment workflows.

You will collaborate across platform engineering, security, applications, and data teams to ensure secure, compliant operations and participate in on-call rotations. The role is on-site in Virginia/Hampton and requires extensive AWS, Kubernetes, Terraform, CI/CD

Qualifications

  • Bachelor's degree and 8–10 years of relevant SRE/DevOps or cloud engineering experience.
  • Expert knowledge of AWS services across compute, networking, storage, IAM, and serverless components.
  • Strong experience with Terraform and CloudFormation and automation principles.
  • Experience building CI/CD pipelines and progressive delivery using GitHub Actions or similar.
  • Deep understanding of Kubernetes administration and Docker deployments.
  • Proven DR validation, system rebuilds, and data integrity checks.
  • Proficiency with Python or Java or Go.
  • Experience debugging complex failures and incident response.
  • Strong analytical, documentation, and communication skills.
  • Ability to work in fast-paced environments.

Responsibilities

  • Support DR drill execution and validate platform rebuild procedures.
  • Execute infrastructure drills to meet 48-hour recovery targets.
  • Verify data completeness and integrity during drills and document findings.
  • Identify readiness gaps in infrastructure, automation, monitoring, and data recovery.
  • Design and support IaC workflows using Terraform, CloudFormation, and CI/CD pipelines.
  • Manage Kubernetes clusters and containerized workloads; scale and improve reliability.
  • Build observability with CloudWatch, Datadog, and other tools.
  • Develop automation scripts in Python or Java to reduce manual work.
  • Collaborate with platform engineering, security, applications, and data teams.
  • Participate in on-call rotations and perform root cause analysis.

Skills

AWS services
CI/CD pipelines
Kubernetes administration
IaC: Terraform/CloudFormation
Programming: Python/Java/C#
Incident response
Monitoring/observability
Data integrity checks

Education

Bachelor's degree in related field

Tools

Terraform
AWS CloudFormation
GitHub Actions
Datadog
CloudWatch
Jenkins
GitLab

Job description

Role Summary

A Senior Site Reliability Engineer is responsible for maintaining the reliability, resilience, and recoverability of mission-critical cloud-based platforms. This role involves leading infrastructure disaster recovery drills, validating platform rebuild procedures, and automating deployment workflows. The engineer collaborates across teams to enhance complex environments, support modernization efforts, and ensure operational continuity.

Responsibilities
  • Support full lifecycle platform portability and disaster recovery (DR) drill execution, including validation of platform rebuild procedures and DR playbooks
  • Execute infrastructure-level drill activities to ensure the platform can be fully rebuilt within the 48-hour recovery target
  • Verify end-to-end data completeness, integrity, and accuracy during drill exercises, documenting results and remediation recommendations
  • Identify exit readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes, and drive corrective actions
  • Design, implement, and support automated Infrastructure as Code (IaC) workflows utilizing Terraform, AWS CloudFormation, and CI/CD pipelines
  • Manage and optimize Kubernetes clusters and containerized workloads, including provisioning, scaling, and reliability enhancements
  • Build and maintain observability solutions using CloudWatch, Datadog, and other monitoring tools for service reliability and proactive incident response
  • Develop automation, tooling, and scripts using Python or Java to reduce manual processes and improve operational repeatability
  • Collaborate with platform engineering, security, applications, and data teams to ensure secure, compliant, and consistent platform operations
  • Participate in on-call rotations, root cause analysis, and incident response activities to strengthen system resilience and operational excellence
Qualifications
  • Bachelor's degree and 8-10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; Master's degree with 6-8 years; or equivalent practical experience
  • Expert knowledge of AWS services across compute, networking, storage, IAM, and serverless components
  • Strong experience with Infrastructure as Code (Terraform, CloudFormation) and automation principles
  • Experience building CI/CD deployment pipelines and implementing progressive delivery mechanisms using GitHub Actions or similar tools
  • Deep understanding of Kubernetes administration, container orchestration, and Docker deployments
  • Proven experience validating disaster recovery processes, performing system rebuilds, and conducting data integrity checksExperience with monitoring tools like CloudWatch, Datadog, or comparable observability solutions
  • Proficiency with programming or scripting languages such as Python, Java, C#, or Go
  • Experience debugging complex failure modes, network issues, backpressure, and consistency challenges
  • Strong analytical, documentation, and communication skills within technical teams
  • Ability to work in a fast-paced environment supporting mission-critical systems
Preferred Qualifications
  • AWS DevSecOps Engineer certification (preferred)
  • Additional AWS certifications (Solutions Architect, SysOps, Developer) and Kubernetes certifications (CKA, CKAD)
  • Familiarity with Zero Trust security models and cloud security best practices
  • Experience with GitLab, Jenkins, or similar CI/CD platforms
  • Experience working in highly regulated environments, such as healthcare, finance, DHS, DoD, or CMS
  • Background supporting federal, defense, or large enterprise programs involving legacy-to-cloud modernization
  • Prior involvement in large-scale disaster recovery drills, continuity of operations (COOP), or portability/executable readiness assessments

Publishing Pay Range: $68.00- $72.00 Hourly

This position is based in office and requires employee to work on-site.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Peraton • United States

Remote
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Good co India • United States

Remote
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

VantageScore® • San Francisco (CA)

On-site
USD 135,000 - 165,000
Medical insurance
Dental insurance
401(k) plan
+1
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Pentangle Tech Services | P5 Group • Lowell (MA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

On-site
USD 150,000 - 155,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

On-site
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Cloud Site Reliability Engineer (SRE)
Senior Cloud Site Reliability Engineer (SRE)

Peraton • Washington

Remote
USD 130,000 - 170,000
Remote work