AWS Cloud Platforms Site Reliability Engineer

ECS

Fairfax (VA)

Hybrid

USD 140,000 - 180,000

Full time

13 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Everforth ECS in Fairfax, VA, is seeking an experienced AWS Cloud Platforms Site Reliability Engineer to join in a hybrid capacity. The role focuses on reliability, resiliency, and automation to design, operate, and improve an Azure Government cloud infrastructure supporting DoD workloads.

Candidates should have strong cloud architecture skills and hands-on engineering experience, with Terraform as a primary tool.

Qualifications

  • Experience designing reliable cloud infrastructure for mission-critical workloads.
  • Proven ability to implement HA and DR solutions and automate infrastructure.
  • Security-conscious with DoD and government cloud environments.

Responsibilities

  • Design, deploy, and maintain highly available, fault-tolerant cloud infrastructure using IaC.
  • Support and maintain HA and DR solutions with automated failover and backups.
  • Manage lifecycle of cloud infrastructure including provisioning, patching, and decommissioning.
  • Implement monitoring, logging, alerting, and observability to detect issues before impact.
  • Participate in incident response and root-cause analyses to prevent recurrence.

Skills

Cloud architecture
Reliability engineering
Terraform
Infrastructure as Code
HA/DR
Monitoring
Incident response

Education

Bachelor's degree +5+ years
High School Diploma +9+ years

Tools

Terraform

Job description

Job Description

Everforth ECS is seeking an experienced

Everforth ECS is seeking an experienced AWS Cloud Platforms Site Reliability Engineer to work in our Fairfax, VA office in a hybrid capacity. Everforth ECS is seeking an experienced AWS Cloud Platforms Engineer specializing in reliability and resiliency to design, operate, and continuously improve an Azure Government cloud infrastructure supporting mission-critical workloads for multiple coalition Mission Partner Network enclaves in support of the DoW community. This position will have a strong focus on infrastructure reliability, automation, High Availability (HA), Disaster Recovery (DR), and Infrastructure as Code (IaC), with Terraform serving as a primary platform for provisioning and managing cloud resources. The Cloud Platforms Engineer will work across cloud infrastructure, platform engineering, security, and application teams to ensure production environments remain reliable, scalable, recoverable, secure, and operationally sustainable. The ideal candidate combines a strong cloud architecture skill set with hands-on operational experience and an automation-first approach to infrastructure management.

  • Design, deploy, and maintain highly available, fault-tolerant cloud infrastructure using Terraform and Infrastructure as Code principles.
  • Support and maintain High Availability (HA) and Disaster Recovery (DR) solutions, including infrastructure redundancy, automated failover, backup and restoration, geographic resiliency, and recovery procedures.
  • Manage the complete lifecycle of cloud infrastructure, including provisioning, configuration, operating system and platform maintenance, patching, upgrades, vulnerability remediation, and decommissioning.
  • Support automation for infrastructure deployment and routine operational activities to reduce manual administration, configuration drift, and the potential for human error.
  • Implement and maintain monitoring, logging, alerting, and observability capabilities to identify infrastructure degradation, capacity constraints, performance issues, and potential service disruptions before they impact users.
  • Participate in incident response, troubleshooting, and root-cause analysis for production infrastructure events and develop corrective actions to prevent recurrence.
  • Define and maintain infrastructure reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability targets, recovery time objectives (RTOs), and recovery point objectives (RPOs).
  • Develop, maintain, and regularly validate disaster recovery procedures through recovery exercises, failover testing, and infrastructure restoration testing.
  • Evaluate cloud infrastructure capacity, performance, availability, and scalability and recommend architectural or operational improvements.
  • Partner with application development, cybersecurity, DevOps, and platform engineering teams to establish standardized deployment patterns and resilient cloud architectures.
  • Maintain infrastructure documentation, operational procedures, architecture diagrams, runbooks, and recovery procedures required to support production environments.
  • Provide technical leadership and guidance regarding cloud infrastructure reliability, resiliency, automation, and operational best practices.
  • Other duties, as assigned.

Note: Salary is commensurate with skillset, qualifications, experience, and educational background.

Salary Range: $140,000-180,000

General Description Of Benefits
Required Skills
  • U.S. Citizen.
  • Active DoD Secret security clearance.
  • High School Diploma and 9+ years of relevant experience. Alternatively, a Bachelors in a related field of study and 5+ years of experience.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AWS Cloud Platforms Site Reliability Engineer
AWS Cloud Platforms Site Reliability Engineer

ecsfederal • Virginia (MN)

Hybrid
USD 140,000 - 180,000
Cloud Platforms Engineer
Cloud Platforms Engineer

ECS • Fairfax (VA)

On-site
USD 180,000 - 210,000
AWS Cloud Platforms Site Reliability Engineer
AWS Cloud Platforms Site Reliability Engineer

Everforth ECS • Merrifield (VA)

On-site
USD 120,000 - 160,000
Senior Cloud Engineer
Senior Cloud Engineer

ecsfederal • Virginia (MN)

Hybrid
USD 170,000 - 200,000
Senior Cloud Engineer
Senior Cloud Engineer

ECS • Fairfax (VA)

On-site
USD 170,000 - 200,000
Hybrid work model
AWS Cloud Platforms SRE: Reliability, IaC & HA/DR (Hybrid)
AWS Cloud Platforms SRE: Reliability, IaC & HA/DR (Hybrid)

ecsfederal • Virginia (MN)

Hybrid
USD 140,000 - 180,000
DevOps Engineer
DevOps Engineer

ecsfederal • Maryland

Hybrid
USD 120,000 - 160,000
Cloud Engineer
Cloud Engineer

ECS • Fairfax (VA)

On-site
USD 120,000 - 160,000
Lead Technical Engineer
Lead Technical Engineer

ECS • Fairfax (VA)

Hybrid
USD 175,000 - 210,000
Hybrid AWS Cloud SRE for Resilient Gov Infra
Hybrid AWS Cloud SRE for Resilient Gov Infra

ECS • Fairfax (VA)

Hybrid
USD 140,000 - 180,000