Cloud SRE: Self-Healing, IaC, Kubernetes & GovCloud

ECS

Arlington (TX)

Hybrid

USD 130,000 - 180,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Everforth ECS in Arlington/remote is seeking a Cloud Site Reliability Engineer (SRE) to own reliability for production systems on AWS GovCloud, IL5 zero-trust, with an IaC foundation using Terraform and Kubernetes. The role emphasizes self-healing automation, data-driven SLOs, and scalable infrastructure.

You will define reliability metrics, implement CI/CD, and operate across observability platforms while guiding teams toward automation and government compliance requirements.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, or related field (or equivalent practical experience)
  • 5+ years of SRE experience (or equivalent), with demonstrated technical leadership
  • 10 years of general work experience
  • Track record building self-healing/auto-remediating systems, not just dashboards
  • Jenkins experience
  • Expert AWS knowledge, GovCloud experience strongly preferred
  • Deep Kubernetes and Terraform expertise at production scale
  • Strong software engineering background (Python and/or Go)
  • Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc.)
  • Proven incident command and postmortem experience
  • Strong communication skills across technical and federal leadership audiences
  • Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable)
  • US Citizenship

Responsibilities

  • Self-Healing Operations
  • Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention
  • Shift the team’s posture from “is it running, how do we fix it” to “how do we make it fix itself”
  • Automate service restoration first; investigate root cause after
  • Uptime Goals & Reliability
  • Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them
  • Use live metrics to decide what’s “reliable enough” and where to invest next
  • Infrastructure
  • Enforce infrastructure-as-code and configuration-as-code, no manual tech change
  • Own Terraform standards and reusable modules adopted across programs
  • Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling)
  • Set CI/CD and pipeline-as-code standards, including progressive delivery
  • Observability & Incidents
  • Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct
  • Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation
  • Collaboration & Leadership
  • Partner with development and contractor teams leads to embed reliability and automation across the software
  • Mentor engineers toward this same automation-first philosophy
  • Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation

Skills

SRE leadership
Automation
Self-healing systems
Communication
Incident management

Education

Bachelor’s degree in CS/IT or equivalent

Tools

Terraform
Kubernetes
AWS GovCloud
Grafana
Splunk
Prometheus
Loki
Python
Go

Job description

Everforth ECS in Arlington/remote is seeking a Cloud Site Reliability Engineer (SRE) to own reliability for production systems on AWS GovCloud, IL5 zero-trust, with an IaC foundation using Terraform and Kubernetes. The role emphasizes self-healing automation, data-driven SLOs, and scalable infrastructure.

You will define reliability metrics, implement CI/CD, and operate across observability platforms while guiding teams toward automation and government compliance requirements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Cloud SRE: Self-Healing, Kubernetes & IaC
Remote Cloud SRE: Self-Healing, Kubernetes & IaC

ecsfederal • Virginia (MN)

Hybrid
USD 130,000 - 180,000
Self-Healing Cloud SRE - AWS GovCloud, Kubernetes, Remote
Self-Healing Cloud SRE - AWS GovCloud, Kubernetes, Remote

ECS • Arlington (VA)

Hybrid
USD 130,000 - 180,000
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

ecsfederal • Virginia (MN)

Hybrid
USD 130,000 - 180,000
Hybrid AWS Cloud SRE for Resilient Gov Infra
Hybrid AWS Cloud SRE for Resilient Gov Infra

ECS • Fairfax (VA)

Hybrid
USD 140,000 - 180,000
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

ECS • Arlington (TX)

Hybrid
USD 130,000 - 180,000
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

ECS • Arlington (VA)

Hybrid
USD 130,000 - 180,000
AWS Cloud Platforms SRE: Reliability, IaC & HA/DR (Hybrid)
AWS Cloud Platforms SRE: Reliability, IaC & HA/DR (Hybrid)

ecsfederal • Virginia (MN)

Hybrid
USD 140,000 - 180,000
Staff GovCloud SRE - Platform Reliability Leader
Staff GovCloud SRE - Platform Reliability Leader

Medallia • McLean (VA)

Hybrid
USD 159,000 - 230,000
Health and wellness benefits
401(k)
Paid parental leave
+2
Senior SRE - GovCloud, Kubernetes & Secure Cloud
Senior SRE - GovCloud, Kubernetes & Secure Cloud

Medallia • McLean (VA)

Hybrid
USD 129,000 - 190,000
Medical
Dental
Vision
+3
Senior SRE: Hybrid Cloud Reliability & Automation Leader
Senior SRE: Hybrid Cloud Reliability & Automation Leader

GovCIO • Arlington (VA)

Hybrid
USD 210,000 - 230,000