Site Reliability Engineer

Stelvio Inc.

Town of Texas (WI)

On-site

USD 125,000 - 145,000

Full time

45 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Stelvio Inc. in Wisconsin seeks a Senior Site Reliability Engineer to own reliability for mission-critical, cloud-native platforms operating 24/7.

You will manage high-volume real-time transactions and collaborate with software and infrastructure teams to build resilient services. You’ll lead incident response, define SLOs/SLIs, conduct blameless post‑mortems, and develop automation with Python, Bash or Go.

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, or similar production-focused role.
  • Experience supporting mission-critical, highly available environments.
  • Strong scripting skills: Python, Bash, and/or Go.
  • Solid Linux administration; Windows exposure beneficial.

Responsibilities

  • Monitor availability, reliability, and performance of production systems.
  • Define and track SLOs, SLIs, error budgets, and reliability metrics.
  • Lead incident response and coordinate cross-team diagnosis and resolution.
  • Conduct root cause analysis and post-incident reviews; ensure follow-ups.
  • Build automation to reduce manual ops using Python, Bash or Go.
  • Develop monitoring, logging, tracing, and alerting with Prometheus, Grafana, ELK, Datadog, OpenTelemetry.
  • Collaborate with software, infra, and operations to improve platform reliability.

Job description

Salary: $125,000 – $145,000 USD annually

The Opportunity

We are looking for an experienced Senior Site Reliability Engineer (SRE) to join a growing technology team responsible for highly available, mission-critical platforms operating 24/7.

You’ll work on modern, cloud-native systems that process high volumes of real-time transactions and data. This is a hands‑on senior position where you'll take ownership of system reliability, lead incident response, improve observability and automation, and work closely with engineering and infrastructure teams to build resilient, scalable services.

What You’ll Be Doing
  • Monitor and maintain the availability, reliability and performance of business‑critical production systems.
  • Define and track SLOs, SLIs, error budgets and reliability metrics.
  • Lead production incident response, coordinating technical teams through diagnosis and resolution.
  • Conduct root cause analysis and lead blameless post‑incident reviews, ensuring follow‑up actions are completed.
  • Build automation using Python, Bash or Go to reduce manual operational work.
  • Develop and improve monitoring, logging, tracing and alerting using technologies such as Prometheus, Grafana, OpenTelemetry, ELK and Datadog.
  • Identify performance bottlenecks and troubleshoot issues across both infrastructure and application layers.
  • Support capacity planning and infrastructure scaling for growing transaction volumes.
  • Work with stateful and distributed technologies including SQL databases and event‑streaming platforms such as Kafka or NATS.
  • Collaborate with software engineering, infrastructure and operations teams to improve overall platform reliability.
  • Help design, implement and test disaster recovery and business continuity processes.
  • Improve operational documentation, processes and SRE best practices.
  • Mentor engineers and help establish a strong reliability engineering culture.
What We’re Looking For
  • 5+ years of experience within Site Reliability Engineering, DevOps, Platform Engineering or a similar production‑focused role.
  • Experience supporting mission‑critical, highly available or large‑scale production environments.
  • Strong experience with at least one major cloud platform: AWS, Azure or GCP.
  • Hands‑on experience with Kubernetes and Docker.
  • Strong understanding of observability and monitoring using tools such as Prometheus, Grafana, ELK, Datadog or OpenTelemetry.
  • Strong scripting skills using Python, Bash and/or Go.
  • Solid Linux administration experience, with exposure to Windows environments beneficial.
  • Experience leading production incidents, root cause analysis and post‑incident reviews.
  • Understanding of SLOs, SLIs, error budgets and reliability engineering principles.
  • Familiarity with relational databases such as PostgreSQL or MySQL and broader distributed‑system concepts.
  • Strong troubleshooting and problem‑solving ability.
  • Excellent communication skills and the ability to collaborate across engineering and operational teams.
Nice to Have
  • Terraform, Ansible and/or Helm.
  • GitOps experience using tools such such as Argo CD.
  • CI/CD and pipeline‑as‑code experience, including GitHub Actions or Dagger.
  • Experience with Kafka, NATS or other event‑streaming technologies.
  • Experience supporting high‑volume transactional or payment systems.
  • Exposure to Intelligent Transportation Systems (ITS), tolling, traffic management, sensor processing or other real‑time infrastructure.
  • Experience designing and testing disaster recovery strategies.
Why Join?

This is an opportunity to work on complex, real‑world technology where reliability directly matters. You’ll have significant ownership across production reliability, observability, automation and incident management while helping shape SRE practices as the platform continues to scale.

Salary: $125,000 – $145,000 USD annually, dependent on experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic Technologies • Dunwoody (GA)

On-site
USD 200,000 - 225,000