SRE Platform Engineer

CACI International Inc

Washington (District of Columbia)

On-site

USD 115,000 - 252,000

Full time

22 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

CACI is seeking a seasoned SRE Platform Engineer to support the DHS Office of the Inspector General, ensuring the reliability, performance, and availability of AI and data analytics platforms used for national security oversight.

You will monitor production environments, implement comprehensive observability, lead incident response, optimize performance and costs, and ensure high availability within secure Azure Government environments.

Qualifications

  • Bachelor's degree plus 15 years in SRE/DevOps or related field; equivalents considered.
  • Active DHS/EOD clearance as required.
  • Extensive SRE principles: monitoring, observability, incident response, capacity planning, performance optimization.

Responsibilities

  • Monitor, maintain, and support production and non-production environments to ensure availability and performance.
  • Implement observability with dashboards, health checks, and log analysis to enable proactive incident detection.
  • Lead incident response including root cause analysis and rapid restoration of service.
  • Analyze metrics across applications and data pipelines to identify bottlenecks and optimize resources.
  • Support capacity planning, autoscaling, and cost optimization for compute and storage.
  • Maintain disaster recovery, high availability, and business continuity capabilities.
  • Monitor Azure Databricks clusters, data pipelines, and analytical workloads for failures.
  • Develop runbooks and documentation to support ongoing operations.

Skills

SRE Principles
Azure Cloud
Monitoring & Observability
Incident Response
Analytical Skills

Education

Bachelor's degree + 15 years experience

Tools

Azure Monitor
Application Insights
Grafana
Prometheus
ELK Stack

Job description

Job Title: SRE Platform Engineer


Job Category: Information Technology


Time Type: Full time


Minimum Clearance Required to Start: None


Employee Type: Regular


Percentage of Travel Required: None


Type of Travel: None


* * *


The Opportunity

CACI is seeking a seasoned Site Reliability (SRE) Platform Engineer to support the Department of Homeland Security (DHS) Office of the Inspector General (OIG). This role offers a unique opportunity to ensure the reliability, performance, and availability of cutting-edge AI and data analytics platforms that strengthen national security oversight through investigative, audit, and inspection operations. As an SRE Platform Engineer, you will be the operational guardian of mission-critical systems including OIG Chat—a large language model assistant—enterprise data platforms, and AI-powered applications that federal oversight professionals depend on daily. You will monitor production environments, implement comprehensive alerting and observability, troubleshoot complex incidents, optimize performance and costs, and ensure high availability through proactive capacity planning and reliability engineering. From analyzing performance metrics to identify bottlenecks to supporting disaster recovery and operational readiness, you will apply site reliability engineering principles to maintain service excellence. Working within secure Azure Government environments, you will build operational resilience for a transformational program of national importance. Join us to make a meaningful impact by ensuring mission-critical AI and data capabilities are always available, performant, and reliable.


Responsibilities


  • Monitor, maintain, and support production and non-production environments for OIGChat, AI applications, and enterprise data platforms to ensure availability, performance, service health, and adherence to service level objectives (SLOs)

  • Implement comprehensive observability including alerting, dashboards, health checks, synthetic monitoring, and log analysis using Azure Monitor, Application Insights, Log Analytics, or equivalent tools to enable proactive incident detection

  • Lead incident response and troubleshooting efforts including root cause analysis, defect resolution, dependency updates, integration validation, and coordination of emergency changes to restore service rapidly

  • Analyze performance and usage metrics across applications, APIs, AI model endpoints, data pipelines, and infrastructure to identify and remediate bottlenecks, latency issues, resource constraints, and efficiency opportunities

  • Support capacity planning, resource sizing, autoscaling configuration, and cost optimization for compute, storage, and AI model consumption to balance performance requirements with fiscal responsibility

  • Implement and maintain backup/restore processes, disaster recovery procedures, high availability architectures, and business continuity capabilities to ensure data protection and operational resilience

  • Monitor data platform availability including Azure Databricks clusters, data pipelines, storage services, and analytical workloads with alerting for pipeline failures, job errors, and performance degradation

  • Support deployment and operational readiness for pilot applications and new capabilities including pre-production validation, performance testing, runbook development, and go-live coordination

  • Provide surge support for complex technical issues, large-scale data collection analysis, analytical environment optimization, and specialized troubleshooting requiring deep platform knowledge

  • Develop and maintain operational documentation including runbooks, troubleshooting guides, architecture diagrams, incident post-mortems, and knowledge transfer materials to support sustainable operations


Qualifications

Required –


  • Bachelor's degree + 15 years of experience in site reliability engineering, DevOps, platform engineering, systems administration, or related field; equivalencies considered (Master's + 12 years; 21 years with no degree; AA + 17 years)

  • Must be able to obtain a Active DHS/ EOD Clearance as required.

  • Extensive experience with Site Reliability Engineering (SRE) principles including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering practices

  • Proven expertise with Azure cloud services including compute, storage, networking, monitoring, and platform-as-a-service (PaaS) offerings with deep understanding of operational best practices

  • Strong experience with monitoring and observability tools (Azure Monitor, Application Insights, Grafana, Prometheus, ELK stack) and implementing alerting, dashboards, and log aggregation

  • Demonstrated ability to troubleshoot complex technical issues across application, platform, and infrastructure layers with strong analytical and problem-solving skills


Desired –


  • Experience operating AI/ML platforms, large language model services (Azure OpenAI), data analytics platforms (Databricks, Synapse), or high-scale cloud applications in production environments

  • Hands-on experience with Azure Government or other secure government cloud environments (AWS GovCloud) with understanding of compliance monitoring, security operations, and federal operational requirements

  • Background in federal government, mission-critical systems, or 24/7 operational environments with experience supporting incident response, change management, and operational excellence programs


What You Can Expect

A culture of integrity.

At CACI, we place character and innovation at the center of everything we do. As a valued team member, you’ll be part of a high-performing group dedicated to our customer’s missions and driven by a higher purpose – to ensure the safety of our nation.


An environment of trust.

CACI values the unique contributions that every employee brings to our company and our customers - every day. You’ll have the autonomy to take the time you need through a unique flexible time off benefit and have access to robust learning resources to make your ambitions a reality.


A focus on continuous growth.

Together, we will advance our nation's most critical missions, build on our lengthy track record of business success, and find opportunities to break new ground — in your career and in our legacy.


Pay Range

There are a host of factors that can influence final salary including, but not limited to, geographic location, Federal Government contract labor categories and contract wage rates, relevant prior work experience, specific skills and competencies, education, and certifications. Our employees value the flexibility at CACI that allows them to balance quality work and their personal lives. We offer competitive compensation, benefits and learning and development opportunities. Our broad and competitive mix of benefits options is designed to support and protect employees and their families. At CACI, you will receive comprehensive benefits such as; healthcare, wellness, financial, retirement, family support, continuing education, and time off benefits.


The Proposed Salary Range For This Position Is

$114,600-$252,100


CACI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, age, national origin, disability, status as a protected veteran, or any other protected characteristic.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Platform Engineer
SRE Platform Engineer

CACI International • Washington

On-site
USD 115,000 - 252,000
Lead Engineer - Platform
Lead Engineer - Platform

CACI International Inc • Washington

On-site
USD 115,000 - 252,000
Lead Engineer - Platform
Lead Engineer - Platform

CACI International • Washington

On-site
USD 115,000 - 252,000
Cloud Automation Engineer
Cloud Automation Engineer

CACI International Inc • Washington

On-site
USD 115,000 - 252,000
Security Engineer
Security Engineer

CACI International Inc • Washington

On-site
USD 90,000 - 190,000
Senior Full Stack Developer
Senior Full Stack Developer

CACI International Inc • Washington

On-site
USD 121,000 - 266,000
Senior AI Engineer
Senior AI Engineer

CACI International Inc • Washington

On-site
USD 115,000 - 252,000
AI Engineering Lead
AI Engineering Lead

CACI International • United States

Hybrid
USD 115,000 - 252,000
Chief Data Engineer
Chief Data Engineer

CACI International Inc • Washington

On-site
USD 115,000 - 252,000
AI Engineering Lead
AI Engineering Lead

CACI International • Washington

On-site
USD 115,000 - 252,000
Healthcare benefits
Retirement plan
Continuing education