Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
ServiceTitan is seeking a Senior Site Reliability Engineer to join the Site Reliability & Infrastructure Engineering team. You will help raise reliability by building observability, running on-call, and scaling our cloud-based platform across Azure and AWS.
You will design dashboards, maintain a Kubernetes-based compute platform, and partner with product teams to ensure reliability is built into architectures from day one. Strong SRE fundamentals and collaboration are essential.
Nice-to-have: database experience (not mandatory — databases are monitored by the same team, not owned individually)Strong programming skills with the ability to build web applications — ideally with solid working knowledge of .NET and ASP.NET. We’re also open to strong Python (Flask, FastAPI) or Java (Spring) backgrounds. The coding assessment will be tailored to whichever language/framework you’re most comfortable inExperience with distributed systems and their common failure modes (retries, timeouts, cascading failures)8-10+ years of relevant hands‑on experienceSRE principles: practical experience with SLIs, SLOs, and error budgets — able to speak to how you’ve defined and monitored these on real systems, not just definitionsObservability: deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch) and the ability to translate that understanding across toolsCloud engineering & networking: solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing)Kubernetes (must‑have): strong, hands‑on understanding of Kubernetes as a systemStrong production troubleshooting skills — comfortable diagnosing issues under pressureCI/CD: strong understanding of a CI/CD system — GitHub Actions preferred, but TeamCity, Azure DevOps, or GitLab CI experience is acceptableYou’re someone who enjoys being directly accountable for the reliability of a business‑critical, large‑scale enterprise systemYou feel rewarded by developing an operability culture in a quickly growing and changing environment, and you’re comfortable owning a wide and diverse set of problem areasYou’re comfortable guiding and making decisions with limited information, and capable of operating within the trade‑offs between solving for immediate needs versus bigger‑scale solutionsBeing human isn’t about checking every box on a list. It’s about the experiences we have, people we meet, and the perspectives we share. So, if you have the skills but are hesitant to apply because of your background,