Senior System Reliability Engineer

On-Demand Group

Eagan (MN)

On-site

USD 229,233,000 - 286,541,000

Full time

47 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

On-Demand Group seeks an experienced Senior Site Reliability Engineer to support cloud-native applications in a Digital Commerce environment, with emphasis on AKS on Azure and strong observability. You will partner with DevOps, Platform Engineering and production teams to boost reliability, scalability, performance and security of distributed apps.

The ideal candidate has hands-on Azure, AKS, and observability expertise, with a proactive approach to monitoring, alerting and incident response.

Qualifications

  • 5+ years in software/operations/devops/SRE roles.
  • 2+ years hands-on SRE or cloud-native engineering in Azure.
  • Strong experience with AKS/Kubernetes and observability concepts.
  • Experience monitoring and troubleshooting distributed cloud-native applications.
  • Experience with Git and cross-team collaboration.

Responsibilities

  • Monitor and improve reliability, availability, latency and performance of Kubernetes-based apps on AKS.
  • Design, deploy and operate SRE capabilities for cloud products and services.
  • Maintain monitoring frameworks using Dynatrace, Azure Monitor and Application Insights.
  • Develop proactive monitoring and alerting to detect performance degradation.
  • Analyze telemetry to identify bottlenecks and optimize distributed systems.
  • Troubleshoot complex cloud, app, infra issues and drive improvements.
  • Automate capabilities to increase reliability, scalability, performance and security.
  • Collaborate with DevOps, Platform Engineering, Production Support, Infrastructure, Network, Security and Development teams.
  • Contribute to CI/CD design and operational processes.
  • Document designs, procedures and monitoring standards.

Skills

Azure
AKS / Kubernetes
Observability
Application performance monitoring
DevOps
Troubleshooting
Cloud-native engineering
Git
Distributed systems

Tools

Dynatrace
Azure Monitor
Application Insights

Job description

We are seeking an experienced Senior Site Reliability Engineer (SRE) to support highly available, cloud-native applications within a Digital Commerce environment.

This role will focus heavily on monitoring, observability, reliability, and performance of Kubernetes-based applications and services running within Microsoft Azure and Azure Kubernetes Service (AKS).

The Senior SRE will partner closely with DevOps, Platform Engineering, Production Support, Infrastructure, Network, Security, Architecture, and Development teams to improve the reliability, scalability, performance, and security of distributed applications.

The ideal candidate brings strong hands‑on experience with Azure, AKS/Kubernetes, observability and application performance monitoring, along with the troubleshooting skills necessary to identify and resolve complex production issues.

Key Responsibilities
  • Monitor and improve the reliability, availability, latency, and performance of Kubernetes-based applications and services running on Azure Kubernetes Service (AKS).
  • Plan, design, deploy, and operate Site Reliability Engineering capabilities for cloud-based products and services.
  • Design and maintain monitoring and observability frameworks using tools such as Dynatrace, Azure Monitor, and Application Insights.
  • Develop monitoring and alerting that proactively identifies symptoms and performance degradation rather than simply reporting outages.
  • Analyze telemetry, logs, metrics and monitoring data to identify application and infrastructure bottlenecks.
  • Recognize and address substandard application or infrastructure performance based on established KPIs.
  • Troubleshoot complex issues across distributed, cloud-native systems.
  • Continuously automate and improve capabilities to increase reliability, scalability, performance, and security.
  • Partner with DevOps and Platform Engineering teams to integrate monitoring and observability capabilities into applications, infrastructure and automated pipelines.
  • Work closely with Infrastructure, Network, Security, Architecture and Development teams to build and maintain highly available Azure environments.
  • Support incident response and help identify root causes of production reliability and performance issues.
  • Design and configure proactive alerting mechanisms and thresholds to enable rapid identification and resolution of issues.
  • Contribute to the design and implementation of CI/CD pipelines and automated operational processes.
  • Document processes, technical designs, operational procedures and monitoring standards.
  • Identify cross-team operational risks and drive issues toward resolution through engineering, troubleshooting and operational improvements.
  • Participate in regulatory and compliance activities as needed.
Required Qualifications
  • 5+ years of experience in Software Engineering, Systems Engineering, Operations Engineering, DevOps, SRE, or related technical roles.
  • 2+ years of hands‑on Site Reliability Engineering, DevOps, or similar cloud‑native engineering experience.
  • Strong experience supporting cloud‑native applications hosted within Microsoft Azure.
  • Hands‑on experience supporting and monitoring Kubernetes / Azure Kubernetes Service (AKS) environments.
  • Experience monitoring application availability, uptime, latency, infrastructure and performance across large distributed systems.
  • Strong knowledge of observability and application performance monitoring concepts.
  • Experience troubleshooting complex cloud, application, infrastructure and system‑related issues.
  • Strong debugging and problem‑solving skills within distributed environments.
  • Experience with version control systems such as Git.
  • Working knowledge across systems, networking, security, databases, storage and cloud infrastructure.
  • Experience collaborating across DevOps, Platform Engineering, Production Support, Infrastructure, Architecture and Development teams.
  • Strong written and verbal communication skills with the ability to communicate technical monitoring and reliability insights to both technical and non‑technical stakeholders.
Preferred Qualifications
  • Deep experience monitoring Kubernetes / AKS environments, containerized applications, services and workloads.
  • Hands‑on experience with Dynatrace.
  • Experience with Azure Monitor and Application Insights.
  • Experience designing and implementing enterprise monitoring and observability frameworks.
  • Experience configuring proactive, symptom‑based alerting and thresholds.
  • Experience analyzing telemetry and monitoring data to identify performance bottlenecks.
  • Strong understanding of application performance monitoring within distributed and microservices‑based environments.
  • Experience improving reliability through automation and SRE practices.
  • Experience supporting high‑volume, customer‑facing web or digital commerce applications.
  • Experience participating in incident response, root cause analysis and continuous operational improvement.
What We're Looking For

The strongest candidate will bring a combination of Site Reliability Engineering, Azure, Kubernetes/AKS and observability expertise.

This is not simply a traditional DevOps or cloud infrastructure role. We are looking for someone who understands how to use monitoring and telemetry to determine what is happening across complex distributed applications, proactively identify performance or reliability issues and work across engineering teams to resolve the underlying problems.

Candidates with hands‑on experience using Dynatrace, Azure Monitor, Application Insights and Kubernetes/AKS observability will be particularly relevant.

The projected hourly range for this position is $80–$100.

On‑Demand Group (ODG) provides employee benefits which includes healthcare, dental and vision insurance. ODG is an equal opportunity employer that does not discriminate on the basis of race, color, religion, gender, sexual orientation, age, national origin, disability or any other characteristic protected by law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Sr
Site Reliability Engineer - Sr

Horizontal Talent • Mendota Heights (MN)

On-site
USD 92,000 - 142,000
Medical, dental, vision
Retirement plan
Senior SRE - Azure
Senior SRE - Azure

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 100,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Moultrie • Birmingham (AL)

On-site
USD 110,000 - 170,000
Senior SRE: Azure AKS & Observability Expert
Senior SRE: Azure AKS & Observability Expert

On-Demand Group • Eagan (MN)

On-site
USD 229,233,000 - 286,541,000
Infrastructure/Cloud DevOps - SRE
Infrastructure/Cloud DevOps - SRE

Bayside Solutions • Cupertino (CA)

On-site
USD 150,000 - 230,000
Sr. Site Reliability Engineer(Local to Atlanta GA Only)
Sr. Site Reliability Engineer(Local to Atlanta GA Only)

Trigint Solutions LLC • Atlanta (GA)

Hybrid
USD 124,000 - 220,000
Senior SRE: Azure & Kubernetes Reliability
Senior SRE: Azure & Kubernetes Reliability

Horizontal Talent • Mendota Heights (MN)

On-site
USD 92,000 - 142,000
Medical, dental, vision
Retirement plan
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

On-site
USD 100,000 - 135,000
Site Reliability Engineer
Site Reliability Engineer

OneStream Software LLC • Northern (KY)

On-site
USD 114,000 - 148,000
Site Reliability Engineer
Site Reliability Engineer

Cosm Inc. • El Segundo (CA), Northern (KY)

On-site
USD 110,000 - 145,000