Site Reliability Engineer - Sr

Horizontal Talent

Mendota Heights (MN)

On-site

USD 92,000 - 142,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision
Retirement plan

Job summary

Horizontal Talent is seeking a Senior Site Reliability Engineer to strengthen monitoring, automation, and incident response for Azure-based Kubernetes applications. You will collaborate across DevOps, security, and development teams to improve reliability, performance, and security in a cloud-native environment.

Qualified candidates bring 5+ years in SRE/ops, 2+ years in cloud-native DevOps, and hands-on Azure/AKS expertise.

Qualifications

  • 5+ years in software or operations engineering or a related technical role.
  • 2+ years in DevOps, Site Reliability Engineering, or a similar cloud-native environment.
  • Hands-on experience supporting cloud-based web applications in Microsoft Azure.
  • Strong knowledge of monitoring infrastructure, uptime, latency, and performance in distributed systems.
  • Experience building or improving CI/CD pipelines.

Responsibilities

  • Support the reliability, availability, performance, and efficiency of cloud services and supporting infrastructure.
  • Design, implement, and improve Site Reliability Engineering practices for cloud products and services.
  • Build and enhance monitoring that detects symptoms early and helps prevent service disruptions.
  • Monitor and troubleshoot Kubernetes-based applications and services running in Azure Kubernetes Service (AKS).
  • Collaborate with DevOps, security, architecture, infrastructure, network, and development teams to resolve cross-functional issues.
  • Analyze key performance indicators and telemetry to identify trends, bottlenecks, and areas for improvement.
  • Automate and strengthen operational processes to improve resilience, scalability, and security.
  • Document processes, findings, and supporting materials to improve clarity and team knowledge sharing.
  • Participate in compliance and regulatory activities as needed.

Skills

SRE experience
DevOps experience
Azure knowledge
CI/CD pipelines
Monitoring/observability
Troubleshooting
Networking basics
Version control
Communication skills
Cross-team collaboration

Education

Bachelor’s degree or equivalent

Tools

AKS
Dynatrace
Azure Monitor
Application Insights
Git

Job description

Join a team focused on keeping cloud-based digital commerce services reliable, secure, and high-performing. This Sr. Site Reliability Engineer role offers the opportunity to strengthen monitoring, automation, and incident response across Kubernetes-based applications running in Azure.

Responsibilities
  • Support the reliability, availability, performance, and efficiency of cloud services and supporting infrastructure
  • Design, implement, and improve Site Reliability Engineering practices for cloud products and services
  • Build and enhance monitoring that detects symptoms early and helps prevent service disruptions
  • Monitor and troubleshoot Kubernetes-based applications and services running in Azure Kubernetes Service (AKS)
  • Collaborate with DevOps, security, architecture, infrastructure, network, and development teams to resolve cross-functional issues
  • Analyze key performance indicators and telemetry to identify trends, bottlenecks, and areas for improvement
  • Automate and strengthen operational processes to improve resilience, scalability, and security
  • Document processes, findings, and supporting materials to improve clarity and team knowledge sharing
  • Participate in compliance and regulatory activities as needed
Skills
  • 5+ years of experience in software engineering, operations engineering, or a related technical role
  • 2+ years of experience in DevOps, Site Reliability Engineering, or a similar cloud-native environment
  • Hands-on experience supporting cloud-based web applications in Microsoft Azure
  • Strong knowledge of monitoring infrastructure, application uptime, latency, and performance in distributed systems
  • Experience building or improving CI/CD pipelines
  • Solid troubleshooting skills across cloud and systems environments
  • Working knowledge of systems, storage, networking, security, and databases
  • Experience with version control tools such as Git, SVN, or CVS
  • Excellent written and verbal communication skills
  • Ability to collaborate effectively across technical and non-technical teams
Preferred Skills
  • Experience with observability and monitoring tools in Kubernetes environments, especially AKS
  • Familiarity with monitoring solutions such as Dynatrace, Azure Monitor, and Application Insights
  • Experience creating alerting strategies based on proactive, symptom-driven thresholds
  • Background in performance analysis using telemetry and monitoring data
  • Bachelor’s degree in Computer Science, Management Information Systems, or a related field, or equivalent experience
  • A proactive mindset focused on continuous improvement and service reliability

Horizontal is committed to building an inclusive workplace where different backgrounds, perspectives, and experiences are valued. We encourage qualified candidates to apply and join a team that believes equity, respect, and collaboration lead to better outcomes for everyone.

By applying for this position, you acknowledge and agree that Horizontal Talent may contact you regarding your application using automated technology, including phone calls, SMS/text messages, or email, which may be delivered by our virtual AI recruiter, Alex.

We offer competitive compensation and benefits including medical, dental, vision, and retirement. Applications will be accepted for 4 weeks. The pay range for this role is $67 - $103 per hour based on qualifications and experience.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior System Reliability Engineer
Senior System Reliability Engineer

On-Demand Group • Eagan (MN)

On-site
USD 229,233,000 - 286,541,000
Senior SRE: Azure & Kubernetes Reliability
Senior SRE: Azure & Kubernetes Reliability

Horizontal Talent • Mendota Heights (MN)

On-site
USD 92,000 - 142,000
Medical, dental, vision
Retirement plan
Site Reliability Engineer
Site Reliability Engineer

Moultrie • Birmingham (AL)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

OneStream Software LLC • Northern (KY)

On-site
USD 114,000 - 148,000
Senior SRE - Azure
Senior SRE - Azure

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 100,000 - 140,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

On-site
USD 100,000 - 135,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mikealbert • Cincinnati (OH)

On-site
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Cosm Inc. • El Segundo (CA), Northern (KY)

On-site
USD 110,000 - 145,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Concord Technologies • United States

On-site
USD 120,000 - 160,000
401K plan with 6% company match
Flex-time off
Paid parental leave
+3
Site Reliability Engineer
Site Reliability Engineer

Axle • Frederick (MD)

On-site
USD 140,000 - 155,000
Paid Time Off
401K match
Educational Benefits
+5