DevOps Site Reliability Engineer (SRE)

IT Veterans, LLC

Washington (District of Columbia)

Hybrid

USD 140,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

IT Veterans LLC in Washington, DC is seeking a DevOps Site Reliability Engineer (SRE) to enhance reliability, performance, and uptime of a mission-critical enterprise platform. You will implement robust deployment practices, build monitoring, and automate operations across multiple cloud environments to sustain 99.9% availability for DoD systems.

You will join an on-call team, respond to incidents, develop automation scripts, design dashboards, and collaborate with CSPs to resolve outages

Qualifications

  • TS/SCI clearance is required.
  • Experience applying Site Reliability Engineering (SRE) principles.
  • Hands-on deployment and management of containerized apps (Kubernetes).
  • Experience administering multi-cloud environments (AWS, Azure, GCP).
  • Experience with enterprise monitoring, logging, and alerting.

Responsibilities

  • Monitor health and performance with automated telemetry to sustain 99.9% uptime.
  • Participate in on-call rotation and respond to incidents within one hour.
  • Develop and maintain automation scripts for diagnostics and operations.
  • Design dashboards for real-time platform health and incident status.
  • Coordinate with CSPs during outages to expedite resolution.
  • Continuously assess reliability and recommend improvements for scalability.

Skills

SRE principles
Python scripting
Bash scripting
Incident response

Tools

Kubernetes
AWS
Azure
GCP

Job description

DevOps Site Reliability Engineer (SRE)

Location: Washington, D.C.
Clearance Required: TS/SCI


Position Overview

IT Veterans is seeking a DevOps Site Reliability Engineer (SRE) to support the reliability, performance, and operational stability of a mission-critical enterprise platform. This position plays a vital role in ensuring continuous availability across multi-cloud environments while supporting software deployments, infrastructure monitoring, incident response, and system automation. The ideal candidate will help bridge development and operations by implementing reliable deployment practices, building robust monitoring capabilities, and rapidly responding to production issues to maintain the required 99.9% platform availability for critical Department of Defense (DoD) systems.


Key Responsibilities:


  • Monitor the health and performance of enterprise infrastructure through continuous system monitoring and automated telemetry to support the required 99.9% platform uptime.

  • Participate in an on‑call rotation and respond to major incidents or platform outages within one hour of notification, executing rapid troubleshooting and system stabilization activities.

  • Develop, maintain, and enhance automation scripts and internal tools that streamline diagnostics, health checks, and routine operational tasks.

  • Design and maintain dashboards that provide real‑time visibility into platform health, including uptime, API performance, incident status, and other key operational metrics.

  • Coordinate directly with Cloud Service Providers (CSPs) during infrastructure outages or service disruptions to expedite issue resolution.

  • Continuously assess system reliability, logging, monitoring, and overall architecture, providing recommendations that improve scalability, resiliency, and operational efficiency.


Required Qualifications:


  • TS/SCI security clearance.

  • Strong understanding of Site Reliability Engineering (SRE) principles and best practices.

  • Hands‑on experience deploying and managing containerized applications using Kubernetes.

  • Experience administering and troubleshooting multi‑cloud environments, including Google Cloud Platform (GCP), Microsoft Azure, and Amazon Web Services (AWS).

  • Experience implementing and maintaining enterprise monitoring, logging, and automated alerting solutions.

  • Proficiency with scripting and automation using languages such as Python, Bash, or similar technologies.


Preferred Qualifications:


  • Passion for building and maintaining highly reliable, mission‑critical systems with demanding uptime requirements.

  • Ability to remain composed and methodical while responding to high‑priority production incidents.

  • Strong troubleshooting, root cause analysis, and diagnostic skills with a focus on rapid issue resolution.

  • A continuous improvement mindset with an emphasis on automation and eliminating repetitive operational tasks.

  • Experience supporting secure, cloud‑native environments within government or defense organizations is a plus.


At IT Veterans LLC, we are committed to providing an environment of mutual respect where equal employment opportunities are available to all applicants and teammates without regard to race, color, religion, sex, pregnancy, national origin, age, physical and mental disability, marital status, sexual orientation, gender identity, gender expression, genetic information, military and veteran status, and any other characteristic protected by applicable law. We believe that diversity and inclusion among our teammates is critical to our success.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

DevOps Site Reliability Engineer (SRE)
DevOps Site Reliability Engineer (SRE)

IT Veterans • Washington

On-site
USD 120,000 - 180,000
DoD Cloud SRE - 99.9% Uptime | TS/SCI Clearance
DoD Cloud SRE - 99.9% Uptime | TS/SCI Clearance

IT Veterans, LLC • Washington

Hybrid
USD 140,000 - 190,000
SRE DevOps Engineer - 99.9% Uptime, Multi-Cloud
SRE DevOps Engineer - 99.9% Uptime, Multi-Cloud

IT Veterans • Washington

On-site
USD 120,000 - 180,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

On-site
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

CruitZi • Washington

On-site
USD 140,000 - 180,000
Systems Engineer SRE- TS/SCI + FS Poly
Systems Engineer SRE- TS/SCI + FS Poly

stanleyreid • Chantilly (VA)

On-site
USD 110,000 - 160,000
Site Reliability Engineer - TS/SCI FSP
Site Reliability Engineer - TS/SCI FSP

Cyrad Solutions LLC • Chantilly (VA)

On-site
USD 220,000 - 276,000
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

ECS • Arlington (VA)

Hybrid
USD 130,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

United States Digital Space LLC • United States

On-site
USD 145,000 - 200,000
Remote-first