Site Reliability Engineer, Cloud & Domain Services

Defense Information Systems Agency

Washington

On-site

USD 144,000 - 187,000

Full time

13 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Defense Information Systems Agency (DISA) seeks a Senior Site Reliability Engineer (Cloud & Domain Services) to lead reliability engineering for core cloud platforms and enterprise domain services. You will build cloud-native IaC automation, improve AD/Entra ID tooling, and enhance Azure Virtual Desktop reliability with automated scaling and recovery.

The role focuses on defining SLOs, architecting deep observability (Azure Monitor, Prometheus, Grafana), and collaborating with Transport SRE to

Qualifications

  • U.S. Citizenship required and eligibility for security clearance.
  • Experience leading reliability engineering for cloud platforms and identity services.

Responsibilities

  • Lead SRE for Cloud & Core Domains, automating reliability of cloud platforms and directory services.
  • Develop IaC templates to deploy Azure services and Kubernetes (AKS) clusters.
  • Automate Active Directory administration and identity synchronization, reducing toil.
  • Engineer automated scaling and recovery for Azure Virtual Desktop environments.
  • Define SLIs/SLOs for cloud availability and authentication latency.

Skills

Site Reliability
Cloud Computing
Active Directory
PowerShell
Infrastructure as Code
Observability
Automation
Systems Administration

Tools

Terraform
AKS
Azure Monitor
Prometheus
Grafana

Job description

Join the Defense Information Systems Agency (DISA), where we deliver the secure IT and communications capabilities that keep our Nation’s Warfighters and senior leaders connected anytime, anywhere. Within Command, Control, Communications and Computers (C4) Enterprise Directorate (J-6), we empower our Nation’s senior leaders and the warfighter community with assured enterprise services to win in an evolving global landscape.

Position Overview: The Senior Site Reliability Engineer (Cloud & Domain Services) leads reliability engineering for core cloud platforms and enterprise domain services, applying automation, observability, and performance engineering to ensure highly available, self‑healing systems. The role builds cloud‑native IaC automation, streamlines identity and directory operations, and enhances the reliability of virtual desktop environments through scalable, automated mechanisms. The Senior SRE defines and measures service reliability across authentication, platform availability, and domain operations, architects deep observability frameworks, and partners with transport SRE counterparts to maintain secure, performant connectivity between cloud services and underlying network infrastructure.

Salary: $143,913 to $187,093 per year(inclusive of TLMS and locality adjustments)

Location: Fort Meade, MD; Arlington, VA

Major Duties:

  • Pioneer SRE for Cloud & Core Domains: Lead the cultural SRE transformation within the J6 operations organization, focusing specifically on automating the reliability of cloud platforms, virtual desktops, and directory services.
  • Engineer Cloud Platform Automation: Write clean, modular infrastructure-as-code (IaC) templates (e.g., Terraform) to automate the deployment, scaling, and configuration of Azure services and Kubernetes (AKS) clusters.
  • Automate Core Directory & Identity Services: Author scripts (PowerShell, Python) to automate Active Directory (AD/Entra ID) administration, group policy enforcement, and identity synchronization, removing manual "toil."
  • Optimize Virtual Desktop Reliability: Engineer automated scaling, performance monitoring, and rapid recovery mechanisms for Azure Virtual Desktop (AVD) environments to ensure a seamless end-user experience.
  • Define Application & Domain SLOs: Collaborate with stakeholders to define and monitor Service Level Indicators (SLIs) and Service Level Objectives (SLOs) specifically for cloud service availability, authentication latency, and system uptime.
  • Architect Cloud-Native Observability: Design and deploy comprehensive monitoring, logging, and tracing frameworks (e.g., Azure Monitor, Prometheus/Grafana) to gain deep visibility into platform health and application performance.
  • Partner with Transport SRE (15%): Collaborate closely with your Transport/Infrastructure counterpart to ensure highly reliable, performant, and secure connectivity between your core cloud platform and the underlying transport networks.

Key Skills

  • Site Reliability Engineering
  • Cloud Computing
  • Active Directory
  • PowerShell
  • Infrastructure as Code
  • Observability
  • Automation
  • Systems Administration

Conditions of Employment:

  • Must be a U.S. Citizen.
  • This national security position, which may require access to classified information, requires a favorable suitability review and security clearance as a condition of employment. Failure to maintain security eligibility may result in termination.
  • Incumbent is required to submit a Financial Disclosure Statement, OGE-450.
  • This is a drug testing designated position
  • Incumbent is required to travel away from their normal duty station to other locations within the continental United States (CONUS) and/or other locations outside the continental United States (OCONUS) up to 30 percent of the time.

Security Clearance:

  • Minimum Requirement: Must possess an active Top Secret (TS) security clearance at the time of application.
  • Preferred Requirement: An active Top Secret / SCI (TS/SCI) clearance is highly preferred.
  • Condition of Employment: Candidates holding a TS clearance must be eligible for, and successfully obtain and maintain, SCI access upon hire.

A career with the U.S. government provides employees with a comprehensive benefits package. As a federal employee, you and your family will have access to a range of benefits that are designed to make your federal career very rewarding.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Chief Engineer (Platform & Systems)
Chief Engineer (Platform & Systems)

Defense Information Systems Agency • Washington

On-site
USD 169,000 - 197,000
Comprehensive federal benefits
Senior Platform Engineer (Capability Engineering Branch Chief)
Senior Platform Engineer (Capability Engineering Branch Chief)

Defense Information Systems Agency • Washington

On-site
USD 128,000 - 187,000
Technical Product Manager, DevSecOps Platforms & Systems
Technical Product Manager, DevSecOps Platforms & Systems

Defense Information Systems Agency • Washington

On-site
USD 169,000 - 197,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

IP Secure, LLC • San Antonio (TX)

Hybrid
USD 140,000 - 190,000
Medical
Dental
Vision
+3
Lead Site Reliability Engineer (SRE) with Security Clearance
Lead Site Reliability Engineer (SRE) with Security Clearance

IPSecure, Inc • San Antonio (TX)

Hybrid
USD 120,000 - 180,000
Medical, Dental, Vision
401(k) match
Education reimbursement
+2
DevOps Site Reliability Engineer (SRE)
DevOps Site Reliability Engineer (SRE)

IT Veterans • Washington

On-site
USD 120,000 - 180,000
Systems Engineer – DevSecOps/SRE (w/ active Secret)
Systems Engineer – DevSecOps/SRE (w/ active Secret)

Critical Solutions • Norfolk (VA)

On-site
USD 87,000 - 112,000
Medical coverage
Dental coverage
Vision coverage
+4
K8 platform engineer
K8 platform engineer

Seneca Resources Company, LLC • Washington

On-site
USD 185,000 - 230,000
Performance bonuses
Company-paid training/certifications
Referral bonuses
+1
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

IPSecure, Inc • San Antonio (TX)

Hybrid
USD 140,000 - 190,000
Medical
Dental
Vision
+10
Infrastructure Site Reliability Engineer (w/ active Secret)
Infrastructure Site Reliability Engineer (w/ active Secret)

Critical Solutions • Norfolk (VA)

On-site
USD 87,000 - 112,000
Premium health coverage
Dental and vision insurance
Life Insurance
+3