Cloud Platforms Engineer

ECS

Fairfax (VA)

Hybrid

USD 180,000 - 210,000

Full time

9 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Everforth ECS seeks an experienced Cloud Platforms Engineer in Fairfax, VA, to design, operate, and continuously improve Azure Government cloud infrastructure supporting mission-critical workloads for coalition enclaves.

Focus on reliability, HA, DR, and IaC with Terraform as a primary tool; collaboration across security, DevOps, and platform teams to ensure scalable, secure, and sustainable production environments.

Qualifications

  • U.S. Citizen.
  • Active DoD Secret security clearance.
  • Bachelor’s degree with 5+ years of related work experience.
  • Ability to obtain DoD 8140 IAT Level II Security+ within 60 days of hire.
  • Hybrid capacity in Fairfax, VA (up to 3 days in office).
  • Travel <20% across CONUS/OCONUS sites.
  • Strong experience designing, deploying, and supporting highly available, fault‑tolerant cloud infra.
  • Hands‑on Terraform and IaC provisioning and management.

Responsibilities

  • Design, deploy, and maintain highly available cloud infrastructure using Terraform and IaC principles.
  • Architect and implement HA and DR solutions including automated failover, backups, restoration, and geographic resiliency.
  • Manage the complete lifecycle of cloud infrastructure (provisioning, config, OS/PL maintenance, patches, upgrades, vulnerability remediation, decommissioning).
  • Develop automation for infra deployment and routine operational activities to reduce manual tasks and drift.
  • Implement and maintain monitoring, logging, alerting, and observability to identify degradation and potential disruptions.
  • Participate in incident response, troubleshoot, and perform root‑cause analysis for production events.
  • Define and maintain infrastructure reliability standards (SLIs/SLOs, availability targets, RTOs, RPOs).
  • Develop and validate disaster recovery procedures through exercises and tests.

Skills

Terraform
Cloud architecture
HA/DR design
Automation
Monitoring/Observability
Incident response
Security/compliance

Education

Bachelor’s degree

Tools

Terraform

Job description

Everforth ECS is seeking an experienced Cloud Platforms Engineer specializing in reliability and resiliency to work in our Fairfax, VA office in a hybrid capacity.

Everforth ECS is seeking an experienced Cloud Platforms Engineer specializing in reliability and resiliency to design, operate, and continuously improve an Azure Government cloud infrastructure supporting mission‑critical workloads for multiple coalition Mission Partner Network enclaves in support of the DoW community. This position will have a strong focus on infrastructure reliability, automation, High Availability (HA), Disaster Recovery (DR), and Infrastructure as Code (IaC), with Terraform serving as a primary platform for provisioning and managing cloud resources.

The Cloud Platforms Engineer will work across cloud infrastructure, platform engineering, security, and application teams to ensure production environments remain reliable, scalable, recoverable, secure, and operationally sustainable. The ideal candidate combines a strong cloud architecture skill set with hands‑on operational experience and an automation‑first approach to infrastructure management.

Key Responsibilities
  • Design, deploy, and maintain highly available, fault‑tolerant cloud infrastructure using Terraform and Infrastructure as Code principles.
  • Architect and implement High Availability (HA) and Disaster Recovery (DR) solutions, including infrastructure redundancy, automated failover, backup and restoration, geographic resiliency, and recovery procedures.
  • Manage the complete lifecycle of cloud infrastructure, including provisioning, configuration, operating system and platform maintenance, patching, upgrades, vulnerability remediation, and decommissioning.
  • Develop automation for infrastructure deployment and routine operational activities to reduce manual administration, configuration drift, and the potential for human error.
  • Implement and maintain monitoring, logging, alerting, and observability capabilities to identify infrastructure degradation, capacity constraints, performance issues, and potential service disruptions before they impact users.
  • Participate in incident response, troubleshooting, and root‑cause analysis for production infrastructure events and develop corrective actions to prevent recurrence.
  • Define and maintain infrastructure reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability targets, recovery time objectives (RTOs), and recovery point objectives (RPOs).
  • Develop, maintain, and regularly validate disaster recovery procedures through recovery exercises, failover testing, and infrastructure restoration testing.
  • Evaluate cloud infrastructure capacity, performance, availability, and scalability and recommend architectural or operational improvements.
  • Partner with application development, cybersecurity, DevOps, and platform engineering teams to establish standardized deployment patterns and resilient cloud architectures.
  • Maintain infrastructure documentation, operational procedures, architecture diagrams, runbooks, and recovery procedures required to support production environments.
  • Provide technical leadership and guidance regarding cloud infrastructure reliability, resiliency, automation, and operational best practices.
  • Other duties, as assigned.

Salary Range: $180,000 - $210,000

Benefits

General Description of Benefits

Required Skills
  • U.S. Citizen.
  • Active DoD Secret security clearance.
  • Bachelor’s degree with 5+ years of related work experience.
  • Ability to obtain a DoD 8140 IAT Level II Security+ (or higher) within 60 days of hire.
  • Ability to work in a hybrid capacity in Fairfax, VA (up to 3 days in office).
  • Ability to travel <20% throughout the lifespan of the Program to CONUS / OCONUS customer sites and government installations.
  • Strong experience designing, deploying, and supporting highly available, fault‑tolerant cloud infrastructure.
  • Hands‑on experience with Terraform and Infrastructure as Code (IaC) principles for provisioning and managing cloud resources.
  • Knowledge of High Availability (HA) and Disaster Recovery (DR) architecture, including redundancy, automated failover, backup and restoration, geographic resiliency, and recovery planning.
  • Experience managing the full cloud infrastructure lifecycle, including provisioning, configuration, patching, upgrades, vulnerability remediation, maintenance, and decommissioning.
  • Strong infrastructure automation skills with an emphasis on reducing manual administration, configuration drift, and operational error.
  • Experience implementing and operating monitoring, logging, alerting, and observability solutions for production infrastructure.
  • Strong troubleshooting and diagnostic skills, including incident response, root‑cause analysis, and corrective action development.
  • Understanding of Site Reliability Engineering concepts, including SLIs, SLOs, availability targets, RTOs, and RPOs.
  • Experience developing and validating disaster recovery procedures, including failover exercises, recovery testing, and infrastructure restoration.
  • Ability to evaluate infrastructure capacity, performance, scalability, availability, and resiliency and recommend architectural or operational improvements.
  • Experience working collaboratively with application development, cybersecurity, DevOps, and platform engineering teams.
  • Ability to develop and maintain technical documentation, including architecture diagrams, operational procedures, runbooks, and recovery documentation.
  • Strong understanding of cloud infrastructure security, vulnerability management, and operational best practices.
  • Demonstrated ability to provide technical leadership and guidance in infrastructure reliability, resiliency, automation, and cloud operations.
  • Experience supporting production, enterprise, regulated, or mission‑critical environments is highly desirable.
  • Strong problem‑solving and decision‑making capabilities, with a proven ability to weigh the relative costs and benefits of potential actions and identify the most appropriate solution.
  • Highly developed interpersonal and oral/written communication skills, with the ability to effectively and professionally interact with a diverse set of stakeholders (from peers to end‑users to executive management).
Desired Skills
  • Bachelor’s degree in Computer Science, Information Technology, Computer Networking, Computer Security, or other Science, Technology, Engineering and Mathematics (STEM) discipline.
  • Demonstrated experience administering and engineering cloud infrastructure in production environments.
  • Strong hands‑on experience with Terraform and Infrastructure as Code (IaC) practices.
  • Experience developing and maintaining infrastructure automation for provisioning and operational tasks.
  • Knowledge of High Availability (HA) and Disaster Recovery (DR) architecture and implementation.
  • Experience with monitoring, logging, alerting, and observability solutions.
  • Strong operational troubleshooting, incident response, and root‑cause analysis skills.
  • Experience supporting enterprise‑scale, regulated, or government environments.
  • Familiarity with security, compliance, vulnerability management, and governance requirements in cloud environments.
  • Ability to work effectively across infrastructure, cybersecurity, DevOps, platform engineering, and application teams.
  • Strong understanding of cloud reliability, resiliency, scalability, and operational best practices.

ECS Federal LLC is an equal opportunity employer and does not discriminate or allow discrimination on the basis of any characteristic protected by law. All qualified applicants will receive consideration for employment without regard to disability, status as a protected veteran or any other status protected by applicable federal, state, or local jurisdiction law.

Everforth ECS is the federal segment of Everforth, a $4B global organization with over 10,000 employees. Our nearly 3,500 professionals deliver advanced technology solutions in data and AI, cybersecurity, and enterprise transformation, serving defense, intelligence, and federal civilian agencies.

Our work powers mission‑critical outcomes, strengthens technology partnerships, and creates meaningful opportunities for our people. We are defined by a commitment to excellence in delivery, a culture of innovation, and an environment where talent can thrive and grow.

We Value
  • Attracting and developing top talent and high‑performing teams
  • Fostering a culture that is engaging, accountable, and mission‑driven
Meet the challenge. Make a difference with Everforth ECS!
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Engineer
Cloud Engineer

ECS • Fairfax (VA)

Hybrid
USD 120,000 - 160,000
Technical Lead
Technical Lead

ECS • Fairfax (VA)

Hybrid
USD 170,000 - 210,000
Cloud Engineer - MID
Cloud Engineer - MID

ECS • Fairfax (VA)

On-site
USD 125,000 - 173,000
Cloud Platform Engineer
Cloud Platform Engineer

ECS • Virginia (MN)

Hybrid
USD 180,000 - 200,000
Cloud Platforms Administrator
Cloud Platforms Administrator

ECS • Fairfax (VA)

On-site
USD 160,000 - 190,000
Project Manager
Project Manager

ECS • Fairfax (VA)

Hybrid
USD 120,000 - 140,000
Senior Cloud Engineer
Senior Cloud Engineer

ECS • Fort Meade (MD)

On-site
USD 170,000 - 210,000
Red Hat OpenShift Cloud Engineer
Red Hat OpenShift Cloud Engineer

ECS • Fairfax (VA)

Hybrid
USD 140,000 - 190,000
Sr. Mission Integration Strategist
Sr. Mission Integration Strategist

ECS • Arlington (VA)

On-site
USD 200,000 - 250,000
Cloud Engineer - MID
Cloud Engineer - MID

Socket.dev • Fairfax (VA)

On-site
USD 125,000 - 173,000
Benefits package