Senior SRE: Cloud, Kubernetes & Observability Lead

GigFinder.ai

Indiana (PA)

On-site

USD 138,000 - 207,000

Full time

16 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Flexible time off
Onboarding program
Leadership training
Medical benefits
401k match
Parental leave
Fertility support

Job summary

ServiceTitan is seeking a Senior Site Reliability Engineer to own reliability for its cloud-based platform. You will work across Kubernetes, observability, and CI/CD, shaping SRE practices and incident response at scale.

You'll diagnose production issues, design dashboards, and partner with engineering teams to ensure scalable, reliable systems while reducing manual toil through automation. Strong Kubernetes and cloud experience are essential.

Qualifications

  • Hands-on Kubernetes administration and troubleshooting in production.
  • Experience defining and monitoring SLIs/SLOs with error budgets.
  • Strong cloud networking fundamentals in AWS or Azure.
  • Experience with modern observability stacks and dashboards.
  • CI/CD systems experience, GitHub Actions preferred.
  • Strong programming skills for web apps (.NET/ASP.NET, Python, or Java).
  • Experience with distributed systems and common failure modes.
  • Excellent production troubleshooting under pressure.
  • 8–10+ years of hands-on experience.

Responsibilities

  • Participate in an on-call rotation and diagnose production issues.
  • Design and maintain observability dashboards and alerting using SLIs/SLOs.
  • Operate and improve a Kubernetes-based compute platform.
  • Collaborate across cloud networking (Azure/AWS) for reliability.
  • Lead root-cause analyses and remediation after incidents.
  • Review architecture with product engineering teams before shipping.
  • Automate repetitive operational tasks and write runbooks.
  • Contribute to CI/CD pipelines to ship changes safely.

Skills

Kubernetes
SRE principles
Cloud networking
Observability
CI/CD
Programming (web apps)
Distributed systems
Production troubleshooting
on-call experience

Tools

OpenTelemetry
Prometheus
Grafana
Datadog
Elasticsearch
GitHub Actions
Azure DevOps
GitLab CI

Job description

ServiceTitan is seeking a Senior Site Reliability Engineer to own reliability for its cloud-based platform. You will work across Kubernetes, observability, and CI/CD, shaping SRE practices and incident response at scale.

You'll diagnose production issues, design dashboards, and partner with engineering teams to ensure scalable, reliable systems while reducing manual toil through automation. Strong Kubernetes and cloud experience are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE - Kubernetes, Cloud Reliability & Observability
Senior SRE - Kubernetes, Cloud Reliability & Observability

ServiceTitan, Inc. • California (MO)

On-site
USD 148,000 - 221,000
Flexible time off
Onboarding program
Annual bonus
+5
Senior Cloud SRE: Observability, Kubernetes & Automation
Senior Cloud SRE: Observability, Kubernetes & Automation

Socket.dev • California (MO)

Hybrid
USD 170,000 - 210,000
Flextime & autonomous work
Health & wellness benefits
Parental leave & fertility support
Senior Director, SRE & Cloud Reliability
Senior Director, SRE & Cloud Reliability

ServiceTitan • United States

On-site
USD 247,000 - 396,000
Flexible time off
Health benefits
Parental leave & fertility support
Senior SRE: Cloud, Kubernetes & Automation
Senior SRE: Cloud, Kubernetes & Automation

Socure • Carson City (NV)

On-site
USD 150,000 - 190,000
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1
Senior SRE & Platform Engineer — Observability & Automation
Senior SRE & Platform Engineer — Observability & Automation

Techunting • United States

On-site
USD 120,000 - 150,000
Head of Cloud Infra & SRE Engineering
Head of Cloud Infra & SRE Engineering

ServiceTitan, Inc. • United States

On-site
USD 247,000 - 396,000
Senior SRE Lead: Cloud Reliability & Automation
Senior SRE Lead: Cloud Reliability & Automation

Oracle • Vienna (VA)

On-site
USD 96,000 - 265,000
Medical, dental, vision insurance
401(k) with company match
Paid time off and holidays
+1
Senior SRE - Remote Cloud Platform Reliability & Automation
Senior SRE - Remote Cloud Platform Reliability & Automation

Piper Companies • United States

Remote
USD 130,000 - 180,000
Medical, dental, vision coverage
401(k)
Paid time off
+1
Senior SRE: Incident & Kubernetes Reliability Lead
Senior SRE: Incident & Kubernetes Reliability Lead

Unique System Skills • United States

Remote
USD 140,000 - 190,000