Senior Cloud SRE: Observability, Kubernetes & Automation

Socket.dev

California (MO)

Hybrid

USD 170,000 - 210,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Flextime & autonomous work
Health & wellness benefits
Parental leave & fertility support

Job summary

ServiceTitan is seeking a Senior Site Reliability Engineer to join our SRE & Infrastructure Engineering team. We run entirely on the cloud, and this team owns the reliability and health of the applications running on top of it — designing the signals that tell us when something’s wrong, and building the systems that keep ServiceTitan running better, faster, and cheaper as we scale.

You will participate in on-call, design dashboards, manage the Kubernetes platform, collaborate across cloud

Qualifications

  • Kubernetes system expertise with hands-on experience.
  • SRE principles: SLIs/SLOs and error budgets.
  • Cloud engineering and networking with AWS or Azure.
  • Observability with a modern stack (OpenTelemetry/Prometheus/Grafana/Datadog/Elasticsearch).
  • CI/CD understanding; GitHub Actions or equivalents.
  • Strong programming skills for web apps (.NET/ASP.NET, Python, or Java).
  • Experience with distributed systems and failure modes.
  • Production troubleshooting under pressure.
  • 8–10+ years of relevant hands-on experience.
  • Nice-to-have: database experience.

Responsibilities

  • Participate in on-call rotation with runbooks and playbooks.
  • Design, build, and maintain dashboards and alerts based on SLIs/SLOs.
  • Operate and improve Kubernetes-based compute platform.
  • Work across Azure/AWS networking to support reliable systems.
  • Investigate production incidents and perform root cause analysis.
  • Partner with product engineering to review architecture before shipping.
  • Build and maintain automation to reduce repetitive work.
  • Write runbooks and documentation for on-call knowledge sharing.
  • Help define non-functional requirements like scalability and availability.
  • Collaborate to adopt reliability and observability best practices.
  • Contribute to CI/CD pipelines and safe deployment processes.

Skills

Kubernetes
SRE principles
Cloud networking
Observability
CI/CD
Programming skills
Distributed systems
Troubleshooting under pressure
Experience 8-10y
Database experience

Tools

OpenTelemetry
Prometheus
Grafana
Datadog
Elasticsearch
GitHub Actions
TeamCity
Azure DevOps
GitLab CI

Job description

ServiceTitan is seeking a Senior Site Reliability Engineer to join our SRE & Infrastructure Engineering team. We run entirely on the cloud, and this team owns the reliability and health of the applications running on top of it — designing the signals that tell us when something’s wrong, and building the systems that keep ServiceTitan running better, faster, and cheaper as we scale.

You will participate in on-call, design dashboards, manage the Kubernetes platform, collaborate across cloud

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud, Kubernetes & Observability Lead
Senior SRE: Cloud, Kubernetes & Observability Lead

GigFinder.ai • Indiana (PA)

On-site
USD 138,000 - 207,000
Flexible time off
Onboarding program
Leadership training
+4
Senior SRE - Kubernetes, Cloud Reliability & Observability
Senior SRE - Kubernetes, Cloud Reliability & Observability

ServiceTitan, Inc. • California (MO)

On-site
USD 148,000 - 221,000
Flexible time off
Onboarding program
Annual bonus
+5
Senior Director, SRE & Cloud Reliability
Senior Director, SRE & Cloud Reliability

ServiceTitan • United States

On-site
USD 247,000 - 396,000
Flexible time off
Health benefits
Parental leave & fertility support
Senior Cloud SRE: Secure, Scalable Infra & Automation
Senior Cloud SRE: Secure, Scalable Infra & Automation

Okta • Bellevue (WA)

On-site
USD 147,000 - 202,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion
Senior Site Reliability Engineer - Scalable Cloud & Impact
Senior Site Reliability Engineer - Scalable Cloud & Impact

ServiceTitan • United States

On-site
USD 120,000 - 180,000
Flexible time off
Fully paid medical, dental, and vision
401(k) with company match
+3
Senior Cloud SRE: Automation, Security & Scale
Senior Cloud SRE: Automation, Security & Scale

Okta • San Francisco (CA)

On-site
USD 165,000 - 226,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion
Head of Cloud Infra & SRE Engineering
Head of Cloud Infra & SRE Engineering

ServiceTitan, Inc. • United States

On-site
USD 247,000 - 396,000
Senior SRE: Cloud, Kubernetes & Automation
Senior SRE: Cloud, Kubernetes & Automation

Socure • Carson City (NV)

On-site
USD 150,000 - 190,000
Senior SRE - Cloud & Observability
Senior SRE - Cloud & Observability

Ridgeline • Reno (NV)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Education reimbursement
Wellness reimbursement
+1
Senior SRE & Automation Engineer (AWS, Observability)
Senior SRE & Automation Engineer (AWS, Observability)

Tata Consultancy Services • Englewood Cliffs (NJ)

On-site
USD 110,000 - 125,000