Cloud Site Reliability Engineer (SRE)

ecsfederal

Virginia (MN)

Hybrid

USD 130,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Everforth ECS seeks a Cloud Site Reliability Engineer (SRE) to own reliability for production systems across our federal cloud. You will build automation, define SLOs, and drive a containerized, IaC-first environment in AWS GovCloud with Kubernetes at scale.

The role combines software engineering with operations, focusing on self-healing infrastructures, incident management, and secure, auditable deployment pipelines. Arlington, VA; hybrid/remote options understood.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, or related field.
  • 10+ years of related work experience in SRE/DevOps or similar.
  • Experience building self-healing/auto-remediating systems and strong automation mindset.
  • Proven track record with AWS GovCloud and security compliance (FedRAMP High).
  • Strong software engineering background (Python and/or Go).

Responsibilities

  • Design and implement automated remediation to detect, respond to, and recover from failures without manual intervention.
  • Shift culture toward self-healing by increasing automation coverage across services.
  • Define data-driven SLOs and error budgets for production systems.
  • Enforce infrastructure-as-code and CI/CD standards; own Terraform modules and pipelines.
  • Build observability frameworks with monitoring, logging, tracing, and incident postmortems; drive incident learning and automation.

Skills

AWS GovCloud
Kubernetes
Terraform
Jenkins
Go
Python
Observability
Incident command

Education

Bachelor’s degree in Computer Science or related field

Tools

Terraform
Jenkins
Grafana
Splunk
Prometheus
Loki
Datadog
AWS GovCloud

Job description

Everforth ECS is seeking a Cloud Site Reliability Engineer (SRE) to work in our Arlington, VA office/remotely.

Our Philosophy

We believe the job of an SRE is to engineer the cloud to run itself. That means writing software and automation that lets systems detect and recover from failure on their own, rather than relying on someone to notice an alert and manually fix it. When something breaks, self-healing comes first, deep root-caused debugging happens after service is restored, not instead of it. We’re looking for someone who automates the operational task by default, not documents the runbook for doing it by hand.

About the Role

This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You’ll define what “reliable enough” looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation.

Responsibilities
Self-Healing Operations
  • Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention
  • Shift the team’s posture from “is it running, how do we fix it” to “how do we make it fix itself”
  • Automate service restoration first; investigate root cause after
Uptime Goals & Reliability
  • Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them
  • Use live metrics to decide what’s “reliable enough” and where to invest next
Infrastructure
  • Enforce infrastructure-as-code and configuration-as-code, no manual tech change
  • Own Terraform standards and reusable modules adopted across programs
  • Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling)
  • Set CI/CD and pipeline-as-code standards, including progressive delivery
Observability & Incidents
  • Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct
  • Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation
Collaboration & Leadership
  • Partner with development and contractor teams leads to embed reliability and automation across the software
  • Mentor engineers toward this same automation-first philosophy
  • Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation

Salary Range: $130,000 - $180,000

General Description of Benefit

  • Bachelor’s degree in Computer Science, Information Technology, or related field (or equivalent practical experience)
  • 5+ years of SRE experience (or equivalent), with demonstrated technical leadership
  • 10 years of general work experience
  • Track record building self-healing/auto-remediating systems, not just dashboards
  • Jenkins experience
  • Expert AWS knowledge, GovCloud experience strongly preferred
  • Deep Kubernetes and Terraform expertise at production scale
  • Strong software engineering background (Python and/or Go)
  • Experience operatingobservability platforms (Grafana, Splunk, Prometheus, Loki, etc.)
  • Proven incident command and postmortem experience
  • Strong communication skills across technical and federal leadership audiences
  • Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable)
  • US Citizenship
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

ECS • Arlington (TX)

Hybrid
USD 130,000 - 180,000
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

ECS • Arlington (VA)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
Cloud SRE: Self-Healing, IaC, Kubernetes & GovCloud
Cloud SRE: Self-Healing, IaC, Kubernetes & GovCloud

ECS • Arlington (TX)

Hybrid
USD 130,000 - 180,000
Remote Cloud SRE: Self-Healing, Kubernetes & IaC
Remote Cloud SRE: Self-Healing, Kubernetes & IaC

ecsfederal • Virginia (MN)

Hybrid
USD 130,000 - 180,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

On-site
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Self-Healing Cloud SRE - AWS GovCloud, Kubernetes, Remote
Self-Healing Cloud SRE - AWS GovCloud, Kubernetes, Remote

ECS • Arlington (VA)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Calance • United States

On-site
USD 150,000 - 200,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Staffing Science • Arizona

On-site
USD 180,000 - 240,000
SRE (Site Realiability Engineer)
SRE (Site Realiability Engineer)

STRATIS Cloud Tech Solutions INC • Arkansas

On-site
USD 110,000 - 150,000
Competitive salary
Growth and learning opportunities
Friendly, collaborative team