AWS Cloud Platforms Site Reliability Engineer

ecsfederal

Virginia (MN)

Hybrid

USD 140,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Everforth ECS seeks an experienced AWS Cloud Platforms Site Reliability Engineer for our Fairfax, VA hybrid office. You will design, deploy, and operate a reliable Azure Government cloud infrastructure, focusing on HA, DR, and IaC with Terraform as a primary tool.

The role requires leading reliability efforts, defining SLIs/SLOs, and collaborating across security, DevOps, and platform teams to ensure scalable, secure production environments. Hybrid work up to 3 days in office.

Qualifications

  • Design, deploy, and maintain highly available cloud infrastructure using Terraform and IaC.
  • Support HA and DR solutions including automated failover, backups, and geographic resiliency.
  • Full cloud lifecycle management: provisioning, patching, upgrades, vulnerability remediation, and decommissioning.
  • Automation to reduce manual administration and configuration drift.
  • Monitoring, logging, alerting, and observability for production infrastructure.
  • Incident response, root-cause analysis, and corrective action development.
  • Define and maintain SLIs, SLOs, and recovery objectives.
  • Validate disaster recovery procedures through drills and testing.
  • Collaborate with application, cybersecurity, DevOps, and platform teams.

Responsibilities

  • Design, deploy, and maintain highly available cloud infrastructure.
  • Support HA/DR solutions and automated failover processes.
  • Manage full cloud infrastructure lifecycle: provisioning, patching, upgrades, remediation.
  • Automate deployment and routine ops to reduce manual work and drift.
  • Implement monitoring, logging, and observability for production.
  • Lead incident response and root-cause analysis; implement fixes.
  • Define reliability standards: SLIs, SLOs, RTOs, RPOs.
  • Regularly validate disaster recovery through tests.
  • Collaborate with cross-functional teams for resilient architectures.

Skills

Cloud infrastructure
HA/DR design
IaC automation
Monitoring observability
Incident response
SRE concepts (SLIs/SLOs)
Security best practices

Education

High School Diploma
Bachelor's in related field

Tools

Terraform

Job description

Everforth ECS is seeking an experienced AWS Cloud Platforms Site Reliability Engineer to work in our Fairfax, VA office in a hybrid capacity.

Everforth ECS is seeking an experienced AWS Cloud Platforms Engineer specializing in reliability and resiliency to design, operate, and continuously improve an Azure Government cloud infrastructure supporting mission-critical workloads for multiple coalition Mission Partner Network enclaves in support of the DoW community. This position will have a strong focus on infrastructure reliability, automation, High Availability (HA), Disaster Recovery (DR), and Infrastructure as Code (IaC), with Terraform serving as a primary platform for provisioning and managing cloud resources.

The Cloud Platforms Engineer will work across cloud infrastructure, platform engineering, security, and application teams to ensure production environments remain reliable, scalable, recoverable, secure, and operationally sustainable. The ideal candidate combines a strong cloud architecture skill set with hands‑on operational experience and an automation‑first approach to infrastructure management.

  • Design, deploy, and maintain highly available, fault‑tolerant cloud infrastructure using Terraform and Infrastructure as Code principles.
  • Support and maintain High Availability (HA) and Disaster Recovery (DR) solutions, including infrastructure redundancy, automated failover, backup and restoration, geographic resiliency, and recovery procedures.
  • Manage the complete lifecycle of cloud infrastructure, including provisioning, configuration, operating system and platform maintenance, patching, upgrades, vulnerability remediation, and decommissioning.
  • Support automation for infrastructure deployment and routine operational activities to reduce manual administration, configuration drift, and the potential for human error.
  • Implement and maintain monitoring, logging, alerting, and observability capabilities to identify infrastructure degradation, capacity constraints, performance issues, and potential service disruptions before they impact users.
  • Participate in incident response, troubleshooting, and root‑cause analysis for production infrastructure events and develop corrective actions to prevent recurrence.
  • Define and maintain infrastructure reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability targets, recovery time objectives (RTOs), and recovery point objectives (RPOs).
  • Develop, maintain, and regularly validate disaster recovery procedures through recovery exercises, failover testing, and infrastructure restoration testing.
  • Evaluate cloud infrastructure capacity, performance, availability, and scalability and recommend architectural or operational improvements.
  • Partner with application development, cybersecurity, DevOps, and platform engineering teams to establish standardized deployment patterns and resilient cloud architectures.
  • Maintain infrastructure documentation, operational procedures, architecture diagrams, runbooks, and recovery procedures required to support production environments.
  • Provide technical leadership and guidance regarding cloud infrastructure reliability, resiliency, automation, and operational best practices.
  • Other duties, as assigned.

Note: Salary is commensurate with skillset, qualifications, experience, and educational background.

Salary Range: $140,000-180,000

General Description of Benefits

  • U.S. Citizen.
  • Active DoD Secret security clearance.
  • High School Diploma and 9+ years of relevant experience. A alternatively, a Bachelors in a related field of study and 5+ years of experience.
  • Required Certifications:
    • DoD 8140 IAT Level II Security+ (or higher).
    • AWS Certified Cloud Practitioner
    • AWS Certified Solutions Architect - Associate
    • AWS Certified CloudOps Engineer - Associate
  • Ability to work in a hybrid capacity in Fairfax, VA (up to 3 days in office).
  • Strong experience designing, deploying, and supporting highly available, fault‑tolerant cloud infrastructure.
  • Hands‑on experience with Terraform and Infrastructure as Code (IaC) principles for provisioning and managing cloud resources.
  • Knowledge of High Availability (HA) and Disaster Recovery (DR) architecture, including redundancy, automated failover, backup and restoration, geographic resiliency, and recovery planning.
  • Experience managing the full cloud infrastructure lifecycle, including provisioning, configuration, patching, upgrades, vulnerability remediation, maintenance, and decommissioning.
  • Strong infrastructure automation skills with an emphasis on reducing manual administration, configuration drift, and operational error.
  • Experience implementing and operating monitoring, logging, alerting, and observability solutions for production infrastructure.
  • Strong troubleshooting and diagnostic skills, including incident response, root‑cause analysis, and corrective action development.
  • Understanding of Site Reliability Engineering concepts, including SLIs, SLOs, availability targets, RTOs, and RPOs.
  • Experience developing and validating disaster recovery procedures, including failover exercises, recovery testing, and infrastructure restoration.
  • Ability to evaluate infrastructure capacity, performance, scalability, availability, and resiliency and recommend architectural or operational improvements.
  • Experience working collaboratively with application development, cybersecurity, DevOps, and platform engineering teams.
  • Ability to develop and maintain technical documentation, including architecture diagrams, operational procedures, runbooks, and recovery documentation.
  • Strong understanding of cloud infrastructure security, vulnerability management, and operational best practices.
  • Demonstrated ability to provide technical leadership and guidance in infrastructure reliability, resiliency, automation, and cloud operations.
  • Experience supporting production, enterprise, regulated, or mission‑critical environments is highly desirable.
  • Strong problem‑solving and decision‑making capabilities, with a proven ability to weigh the relative costs and benefits of potential actions and identify the most appropriate solution.
  • Highly developed interpersonal and oral/written communication skills, with the ability to effectively and professionally interact with a diverse set of stakeholders (from peers to end‑users to executive management).
  • Strong problem‑solving and decision‑making capabilities, with a proven ability to weigh the relative costs and benefits of potential actions and identify the most appropriate solution.
  • Highly developed interpersonal and oral/written communication skills, with the ability to effectively and professionally interact with a diverse set of stakeholders (from peers to end‑users to executive management).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AWS Cloud Platforms Site Reliability Engineer
AWS Cloud Platforms Site Reliability Engineer

ECS • Fairfax (VA)

Hybrid
USD 140,000 - 180,000
AWS Cloud Platforms Site Reliability Engineer
AWS Cloud Platforms Site Reliability Engineer

Everforth ECS • Merrifield (VA)

On-site
USD 120,000 - 160,000
Cloud Platforms Engineer
Cloud Platforms Engineer

ECS • Fairfax (VA)

On-site
USD 180,000 - 210,000
Senior Cloud Engineer
Senior Cloud Engineer

ECS • Fairfax (VA)

On-site
USD 170,000 - 200,000
Hybrid work model
Senior Cloud Engineer
Senior Cloud Engineer

ecsfederal • Virginia (MN)

Hybrid
USD 170,000 - 200,000
Lead Technical Engineer
Lead Technical Engineer

ecsfederal • Virginia (MN)

Hybrid
USD 175,000 - 210,000
Hybrid work model
On-site Fairfax, VA
Lead Technical Engineer
Lead Technical Engineer

ECS • Fairfax (VA)

Hybrid
USD 175,000 - 210,000
Senior Cloud Engineer
Senior Cloud Engineer

Everforth ECS • Merrifield (VA)

On-site
USD 120,000 - 160,000
Cloud Engineer
Cloud Engineer

ECS • Fairfax (VA)

On-site
USD 120,000 - 160,000
Cloud/Network Infrastructure Engineer, Senior
Cloud/Network Infrastructure Engineer, Senior

ECS • Arlington (VA)

On-site
USD 123,000 - 184,000