Senior SRE: AI-Driven Kubernetes Reliability Leader

ServiceTitan, Inc.

California (MO)

Hybrid

USD 148,000 - 221,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Flexible time off
Bonus program
Health, dental, vision

Job summary

ServiceTitan, Inc. is seeking a Senior Site Reliability Engineer to join the Site Reliability & Infrastructure Engineering team.

You will help design signals, build scalable systems, and enhance availability across the ServiceTitan cloud, partnering with engineering to review architectures before shipping. You will work on on-call incidents, automate repetitive tasks with AI agents, and contribute to CI/CD and observability efforts.

Qualifications

  • Kubernetes (must-have) hands-on in a system.
  • AI tools/agents experience to diagnose and automate infrastructure work.
  • Agentic depth: autonomous monitoring and action with concrete details.
  • SRE: applying AI to reliability problems, not just coding aids.
  • SLIs/SLOs and error budgets experience on real systems.
  • Cloud engineering with AWS or Azure networking basics.
  • Observability stack experience (OpenTelemetry/Prometheus/Grafana/etc.).
  • CI/CD systems knowledge (GitHub Actions preferred).
  • Strong programming: .NET/ASP.NET; Python/Flask/FastAPI or Java/Spring.
  • Experience with distributed systems and their failure modes.
  • Strong production troubleshooting under pressure.

Responsibilities

  • Participate in an on-call rotation and diagnose production issues using runbooks.
  • Design, build, and maintain observability dashboards and alerting based on SLIs/SLOs.
  • Operate and improve Kubernetes-based compute platform and infrastructure.
  • Collaborate across cloud networking (AWS/Azure) for reliable systems.
  • Investigate incidents, perform root-cause analysis and remediation.
  • Develop AI agents to automate repetitive SRE tasks.
  • Review architecture/infrastructure with product teams before shipping.
  • Write and maintain runbooks and docs for on-call knowledge sharing.
  • Define non-functional requirements (scalability, availability, performance).
  • Promote best practices in reliability and observability across teams.
  • Contribute to CI/CD pipelines and ship changes safely and quickly.

Skills

Kubernetes
AI-native SRE
Agentic depth
SRE application
SRE principles
Cloud engineering & networking
Observability
CI/CD
Programming
Distributed systems
Production troubleshooting
Database experience (nice-to-have)

Job description

ServiceTitan, Inc. is seeking a Senior Site Reliability Engineer to join the Site Reliability & Infrastructure Engineering team.

You will help design signals, build scalable systems, and enhance availability across the ServiceTitan cloud, partnering with engineering to review architectures before shipping. You will work on on-call incidents, automate repetitive tasks with AI agents, and contribute to CI/CD and observability efforts.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE - AI-Powered Cloud Reliability (Remote)
Senior SRE - AI-Powered Cloud Reliability (Remote)

ServiceTitan, Inc. • Northern (KY)

Hybrid
USD 148,000 - 221,000
Senior SRE — Cloud Reliability & Kubernetes Champion
Senior SRE — Cloud Reliability & Kubernetes Champion

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
Senior Director, SRE & Cloud Reliability
Senior Director, SRE & Cloud Reliability

ServiceTitan • United States

On-site
USD 247,000 - 396,000
Flexible time off
Health benefits
Parental leave & fertility support
Senior Site Reliability Engineer - Scalable Cloud & Impact
Senior Site Reliability Engineer - Scalable Cloud & Impact

ServiceTitan • United States

On-site
USD 120,000 - 180,000
Flexible time off
Fully paid medical, dental, and vision
401(k) with company match
+3
Senior SRE Leader: AI-Driven Reliability & Resilience
Senior SRE Leader: AI-Driven Reliability & Resilience

Jobtailor • Arlington (TX)

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
Senior SRE: AI-Driven Infra & Reliability
Senior SRE: AI-Driven Infra & Reliability

Jobless • Ann Arbor (MI)

Hybrid
USD 180,000 - 240,000
Health Care Coverage
Life Insurance
Health Savings Account
+3
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7