Site Reliability Engineer

Stelvio Inc.

Town of Texas (WI)

On-site

USD 125,000 - 145,000

Full time

48 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Stelvio Inc. is seeking an experienced Senior Site Reliability Engineer to own reliability for mission-critical, cloud-native platforms that run 24/7. You will lead incident responses, improve observability, and collaborate with engineering and infrastructure teams to build scalable services.

The role emphasizes ownership of reliability, SLO/SLI tracking, and automation using Python, Bash, or Go. Experience with Kubernetes, Docker, and multiple cloud platforms is required, with strong Linux

Qualifications

  • 5+ years in Site Reliability Engineering, DevOps, or similar
  • Experience managing mission-critical, high-availability environments
  • Hands-on with Kubernetes and Docker
  • Cloud platform experience (AWS/Azure/GCP)
  • Strong observability/monitoring skills (Prometheus, Grafana, ELK, Datadog, OpenTelemetry)
  • SRE principles: SLOs, SLIs, error budgets
  • Scripting in Python, Bash, or Go
  • Linux administration with Windows exposure beneficial
  • Experience leading production incidents and post-incident reviews
  • Familiar with PostgreSQL or MySQL

Responsibilities

  • Monitor and maintain availability and performance of production systems
  • Define and track SLOs, SLIs, and reliability metrics
  • Lead incident response and post-incident reviews
  • Build automation to reduce manual toil (Python/Bash/Go)
  • Improve monitoring/logging/tracing and alerting (Prometheus, Grafana, OpenTelemetry, ELK, Datadog)
  • Troubleshoot across infrastructure and application layers
  • Support capacity planning and scaling
  • Collaborate across engineering and operations to improve reliability
  • Design and test disaster recovery and business continuity

Skills

5+ years SRE/DevOps
Kubernetes
Docker
Cloud platforms (AWS/Azure/GCP)
Observability (Prometheus/Grafana/Open
Python, Bash, Go
PostgreSQL/MySQL
Incident management
SLOs/SLIs

Education

Bachelor's degree in CS or related

Tools

Prometheus
Grafana
ELK
Datadog
OpenTelemetry
Argo CD

Job description

Salary: $125,000 – $145,000 USD annually

The Opportunity

We are looking for an experienced Senior Site Reliability Engineer (SRE) to join a growing technology team responsible for highly available, mission‑critical platforms operating 24/7.

You’ll work on modern, cloud‑native systems that process high volumes of real‑time transactions and data. This is a hands‑on senior position where you'll take ownership of system reliability, lead incident response, improve observability and automation, and work closely with engineering and infrastructure teams to build resilient, scalable services.

What You’ll Be Doing
  • Monitor and maintain the availability, reliability and performance of business‑critical production systems.
  • Define and track SLOs, SLIs, error budgets and reliability metrics.
  • Lead production incident response, coordinating technical teams through diagnosis and resolution.
  • Conduct root cause analysis and lead blameless post‑incident reviews, ensuring follow‑up actions are completed.
  • Build automation using Python, Bash or Go to reduce manual operational work.
  • Develop and improve monitoring, logging, tracing and alerting using technologies such as Prometheus, Grafana, OpenTelemetry, ELK and Datadog.
  • Identify performance bottlenecks and troubleshoot issues across both infrastructure and application layers.
  • Support capacity planning and infrastructure scaling for growing transaction volumes.
  • Work with stateful and distributed technologies including SQL databases and event‑streaming platforms such as Kafka or NATS.
  • Collaborate with software engineering, infrastructure and operations teams to improve overall platform reliability.
  • Help design, implement and test disaster recovery and business continuity processes.
  • Improve operational documentation, processes and SRE best practices.
  • Mentor engineers and help establish a strong reliability engineering culture.
What We’re Looking For
  • 5+ years of experience within Site Reliability Engineering, DevOps, Platform Engineering or a similar production‑focused role.
  • Experience supporting mission‑critical, highly available or large‑scale production environments.
  • Strong experience with at least one major cloud platform: AWS, Azure or GCP.
  • Hands‑on experience with Kubernetes and Docker.
  • Strong understanding of observability and monitoring using tools such as Prometheus, Grafana, ELK, Datadog or OpenTelemetry.
  • Strong scripting skills using Python, Bash and/or Go.
  • Solid Linux administration experience, with exposure to Windows environments beneficial.
  • Experience leading production incidents, root cause analysis and post‑incident reviews.
  • Understanding of SLOs, SLIs, error budgets and reliability engineering principles.
  • Familiarity with relational databases such as PostgreSQL or MySQL and broader distributed‑system concepts.
  • Strong troubleshooting and problem‑solving ability.Excellent communication skills and the ability to collaborate across engineering and operational teams.
Nice to Have
  • Terraform, Ansible and/or Helm.
  • GitOps experience using tools such as Argo CD.
  • CI/CD and pipeline‑as‑code experience, including GitHub Actions or Dagger.
  • Experience with Kafka, NATS or other event‑streaming technologies.
  • Experience supporting high‑volume transactional or payment systems.
  • Exposure to Intelligent Transportation Systems (ITS), tolling, traffic management, sensor processing or other real‑time infrastructure.
  • Experience designing and testing disaster recovery strategies.
Why Join?

This is an opportunity to work on complex, real‑world technology where reliability directly matters. You’ll have significant ownership across production reliability, observability, automation and incident management while helping shape SRE practices as the platform continues to scale.

Salary: $125,000 – $145,000 USD annually, dependent on experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Supio • San Francisco (CA)

On-site
USD 170,000 - 220,000