SRE Platform Engineer

Jobgether SRL

United States

Hybrid

USD 115,000 - 252,000

Full time

33 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Full-time position with hybrid options
Salary range $114,600–$252,100
Healthcare and wellness benefits
Education and professional development
Flexible time-off
AI and cloud infrastructure focus
Secure Azure Government exposure

Job summary

Jobgether SRL is seeking a Senior SRE Platform Engineer based in the United States to ensure reliability, performance, and availability of AI and data platforms in Azure environments. You will own incident response, capacity planning, and disaster recovery for mission-critical systems supporting government operations.

You will work across application, data, and infra teams to optimize latency, autoscale cloud resources, and implement robust monitoring and runbooks for durable operations.

Qualifications

  • Bachelor’s degree and 15 years in SRE/DevOps/platform engineering.
  • Active DHS/EOD clearance as required.
  • Expert in Azure cloud services and observability.

Responsibilities

  • Monitor, maintain, and support production and non-production environments for availability and performance.
  • Implement comprehensive observability with alerts, dashboards, health checks, and log analysis using Azure tools.
  • Lead incident response and troubleshooting including root cause analysis and change coordination.
  • Analyze metrics to identify bottlenecks and opportunities for improved efficiency.
  • Support capacity planning, autoscaling, and cost optimization across compute and storage.
  • Maintain backups, DR procedures, and high-availability architectures.
  • Support deployment readiness through validation, runbooks, and go-live coordination.
  • Provide surge support for complex issues and large-scale data platforms.
  • Develop and maintain runbooks, diagrams, and post-mortems for operations.

Skills

SRE principles
Monitoring
Incident response
Capacity planning
Performance optimization
Reliability engineering
Azure
Observability
Troubleshooting

Education

Bachelor’s degree
Master’s degree

Tools

Azure Monitor
Application Insights
Grafana
Prometheus
ELK
Databricks
Synapse
Azure OpenAI

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Platform Engineer based in United States. This role focuses on ensuring the reliability, performance, and availability of mission-critical AI and data platforms supporting federal oversight operations. You will serve as an operational guardian for AI assistants, enterprise data platforms, and AI-powered applications used in demanding environments. The position combines observability, incident response, performance engineering, capacity planning, disaster recovery, and cloud operations. Working primarily within secure Azure Government environments, you will help build resilient and highly available platforms. You will collaborate across application, data, infrastructure, and operational teams to troubleshoot complex issues and improve service quality. This is an opportunity to apply advanced SRE practices to large-scale technology supporting important government missions.

Accountabilities
  • Monitor, maintain, and support production and non-production environments, ensuring availability, performance, service health, and adherence to service level objectives.
  • Implement comprehensive observability through alerting, dashboards, health checks, synthetic monitoring, and log analysis using tools such as Azure Monitor, Application Insights, and Log Analytics.
  • Lead incident response and troubleshooting, including root cause analysis, defect resolution, dependency updates, integration validation, and emergency change coordination.
  • Analyze application, API, AI model, data pipeline, and infrastructure metrics to identify bottlenecks, latency, resource constraints, and opportunities for improved efficiency.
  • Support capacity planning, resource sizing, autoscaling, and cost optimization across compute, storage, and AI model consumption.
  • Maintain backup and restore processes, disaster recovery procedures, high-availability architectures, and business continuity capabilities.
  • Monitor data platforms, including Azure Databricks clusters, data pipelines, storage services, and analytical workloads, with appropriate alerting for failures and performance degradation.
  • Support the operational readiness and deployment of new applications and capabilities through pre-production validation, performance testing, runbook creation, and go-live coordination.
  • Provide specialized troubleshooting and surge support for complex technical issues, large-scale data collection and analysis, and analytical environment optimization.
  • Develop and maintain runbooks, troubleshooting guides, architecture diagrams, incident post-mortems, and knowledge-transfer documentation to support sustainable operations.
Requirements
  • Bachelor’s degree plus 15 years of relevant experience in SRE, DevOps, platform engineering, systems administration, or a related field; equivalent combinations include a Master’s degree plus 12 years, 21 years without a degree, or an Associate degree plus 17 years.
  • Ability to obtain an active DHS/EOD clearance as required.
  • Extensive knowledge of SRE principles, including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering.
  • Strong expertise with Azure cloud services covering compute, storage, networking, monitoring, and PaaS offerings, along with strong operational best practices.
  • Hands-on experience with observability and monitoring technologies such as Azure Monitor, Application Insights, Grafana, Prometheus, and ELK.
  • Proven ability to troubleshoot complex issues across application, platform, and infrastructure layers, supported by strong analytical and problem-solving skills.
  • Experience operating AI/ML platforms, Azure OpenAI or other large language model services, Databricks, Synapse, or high-scale cloud applications is highly desirable.
  • Experience with Azure Government or other secure government cloud environments, such as AWS GovCloud, including compliance monitoring and security operations, is an advantage.
  • Background supporting federal government, mission-critical, or 24/7 operational environments, including incident response, change management, and operational excellence programs, is preferred.
Benefits
  • Full-time position with hybrid work options and remote eligibility across the United States.
  • Proposed national salary range of $114,600–$252,100, with final compensation influenced by location, contract requirements, experience, skills, education, and certifications.
  • Comprehensive healthcare and wellness benefits.
  • Financial and retirement benefits.
  • Family support programs.
  • Flexible time-off benefits designed to support work-life balance.
  • Continuing education, learning, and professional development opportunities.
  • Opportunity to work with AI, data analytics, cloud infrastructure, observability, and reliability engineering at significant scale.
  • Exposure to secure Azure Government environments, AI platforms, enterprise data systems, and mission-critical operational challenges.
  • Collaborative environment with opportunities to deepen technical expertise and contribute to continuous platform improvement.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Platform Engineer - AI, Data & Azure Gov
SRE Platform Engineer - AI, Data & Azure Gov

CACI International • Washington

On-site
USD 115,000 - 252,000
SRE/Platform Engineer
SRE/Platform Engineer

Stash Talent Services • Virginia (MN)

On-site
USD 96,432 - 117,096
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Free meals and snacks
+2
Mission-Critical SRE Platform Engineer (Azure AI)
Mission-Critical SRE Platform Engineer (Azure AI)

CACI International • United States

Remote
USD 115,000 - 252,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
SRE Platform Engineer: Azure Gov AI & DataOps
SRE Platform Engineer: Azure Gov AI & DataOps

Jobgether SRL • United States

Hybrid
USD 115,000 - 252,000
Full-time position with hybrid options
Salary range $114,600–$252,100
Healthcare and wellness benefits
+4
SRE Platform Engineer
SRE Platform Engineer

CACI International • Washington

On-site
USD 115,000 - 252,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
SRE Platform Engineer for AI & Data Platforms
SRE Platform Engineer for AI & Data Platforms

CACI International Inc • Washington

On-site
USD 115,000 - 252,000
Senior Forward Deployed Engineer, Public Sector
Senior Forward Deployed Engineer, Public Sector

Sitreps • Washington, San Francisco (CA)

On-site
USD 167,000 - 203,000