Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan

United States

On-site

USD 140,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
401(k) match
Fertility coverage up to $20,000
Paid parental leave: 30 weeks birthing
Adoption support up to $20,000
Bereavement leave
New parent programs
Personal development programs

Job summary

ServiceTitan is seeking a Senior Site Reliability Engineer to join the Site Reliability & Infrastructure Engineering team. You will help raise reliability by building observability, running on-call, and scaling our cloud-based platform across Azure and AWS.

You will design dashboards, maintain a Kubernetes-based compute platform, and partner with product teams to ensure reliability is built into architectures from day one. Strong SRE fundamentals and collaboration are essential.

Qualifications

  • 8-10+ years of hands-on experience in SRE or related fields.
  • Strong programming skills in .NET/ASP.NET, Python or Java.
  • Experience with distributed systems and failure modes (retries, timeouts, cascading failures).
  • Deep observability experience with OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch.
  • AWS or Azure cloud engineering & networking fundamentals.

Responsibilities

  • Participate in an on-call rotation and diagnose production issues using runbooks and playbooks.
  • Design, build, and maintain observability dashboards and alerting based on SLIs/SLOs.
  • Operate and improve a Kubernetes-based compute platform across cloud environments.
  • Investigate incidents, perform root-cause analysis, and follow-up remediation.
  • Collaborate with product engineering to review architecture and infra decisions before ship.
  • Build automation to reduce manual operational work and maintain runbooks/documentation.

Job description

  • We’re looking for a Senior Site Reliability Engineer to join our Site Reliability & Infrastructure Engineering team
  • We run entirely on the cloud, and this team owns the reliability and health of the applications running on top of it — designing the signals that tell us when something’s wrong, and building the systems that keep ServiceTitan running better, faster, and cheaper as we scale
  • Participate in an on‑call rotation, using runbooks and playbooks to diagnose and resolve production issues (e.g., adjusting Horizontal Pod Autoscaler rules in response to load)
  • Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
  • Operate and improve our Kubernetes‑based compute platform, which runs the large majority of our infrastructure
  • Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems
  • Investigate and resolve production incidents, including root‑cause analysis and follow‑up remediation work
  • Partner with product engineering teams to review architecture and infrastructure decisions before they ship
  • Build and maintain automation that reduces manual, repetitive operational work across the team
  • Write and maintain runbooks and documentation so on‑call knowledge is shared across the team, not siloed with one person
  • Help define non‑functional requirements — scalability, availability, performance — for new systems as they’re designed
  • Collaborate across engineering teams to adopt best practices in reliability and observability
  • Contribute to CI/CD pipelines and help teams ship changes safely and quickly
Benefits
  • Flexible time off
  • Fully paid individual medical, dental, and vision coverage
  • Access to on‑demand services like Ginger, Headspace, One Medical, Galileo, Aaptiv, and more
  • Tax free savings via HSA, FSA, and Dependent Care FSA
  • 401(k) match
  • Medical plan with up to $20,000 in fertility coverage
  • Up to 30 weeks of paid leave for birthing parents and up to 16 for non-birthing parents
  • Adoption support including reimbursement of up to $20,000
  • Bereavement leave for loss of a pregnancy or loved one
  • Access to programs for new parents including Maven Maternity, Ready, Set, Food! Wellthy, SNOO Smart Sleeper, and more
  • Personal development programs

Nice-to-have: database experience (not mandatory — databases are monitored by the same team, not owned individually)Strong programming skills with the ability to build web applications — ideally with solid working knowledge of .NET and ASP.NET. We’re also open to strong Python (Flask, FastAPI) or Java (Spring) backgrounds. The coding assessment will be tailored to whichever language/framework you’re most comfortable inExperience with distributed systems and their common failure modes (retries, timeouts, cascading failures)8-10+ years of relevant hands‑on experienceSRE principles: practical experience with SLIs, SLOs, and error budgets — able to speak to how you’ve defined and monitored these on real systems, not just definitionsObservability: deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch) and the ability to translate that understanding across toolsCloud engineering & networking: solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing)Kubernetes (must‑have): strong, hands‑on understanding of Kubernetes as a systemStrong production troubleshooting skills — comfortable diagnosing issues under pressureCI/CD: strong understanding of a CI/CD system — GitHub Actions preferred, but TeamCity, Azure DevOps, or GitLab CI experience is acceptableYou’re someone who enjoys being directly accountable for the reliability of a business‑critical, large‑scale enterprise systemYou feel rewarded by developing an operability culture in a quickly growing and changing environment, and you’re comfortable owning a wide and diverse set of problem areasYou’re comfortable guiding and making decisions with limited information, and capable of operating within the trade‑offs between solving for immediate needs versus bigger‑scale solutionsBeing human isn’t about checking every box on a list. It’s about the experiences we have, people we meet, and the perspectives we share. So, if you have the skills but are hesitant to apply because of your background,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan, Inc. • California (MO)

Hybrid
USD 148,000 - 221,000
Flexible time off
Bonus program
Health, dental, vision
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Free meals and snacks
+2
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

FORT • United States

Hybrid
USD 150,000 - 180,000
Healthcare benefits
Flexible work environment
Large-scale cloud platform project
+1
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hard Rock Digital • United States

Hybrid
USD 150,000 - 210,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000