Senior Site Reliability Engineer (Critical Incident Response)

Liftlab, Inc.

Northern (KY)

Hybrid

USD 140,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

LiftLab, Inc. is seeking a Senior Site Reliability Engineer to join the Critical Incident Response team, leading rapid response to production-critical incidents across the multi-tenant platform.

The role collaborates with the Critical Incident Manager and Principal Data Engineer to drive incidents to resolution, implement postmortems, and strengthen platform reliability. Strong SQL/Python and cloud experience (AWS/Azure) are required, with US-hours coverage and mentorship responsibilities.

Qualifications

  • 6-9 years in SRE, production support, or platform engineering, including incident management.
  • Strong SQL and Python; experience with monitoring/alerting tools and cloud platforms (AWS/Azure).
  • Proven experience leading production incident response in a SaaS or multi-tenant environment.

Responsibilities

  • Lead the technical response to P0/P1 production incidents, coordinating investigation and resolution across teams.
  • Make time-critical decisions (rollback vs. forward-fix) and drive incidents to resolution within SLA.
  • Own root-cause analysis and postmortems, and assign and track follow-up action items.
  • Build and maintain monitoring, alerting, and on-call processes to detect and prevent incidents.
  • Author and improve incident runbooks, escalation paths, and severity classification.
  • Provide US-hours coverage for rapid response and mentor junior incident engineers.

Skills

SQL
Python
Incident management
Monitoring & alerting

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

AWS
Azure
Monitoring tools

Job description

Senior Site Reliability Engineer (Critical Incident Response)

Senior engineer on the P0 / critical incident team, responsible for leading rapid response to production-critical (P0/P1) incidents across LiftLab's multi-tenant platform. Works directly with the Critical Incident Manager to drive incidents to resolution, minimize client impact, and strengthen the reliability of the platform.


6-9 yrs exp USA Remote Full-time Reports to: Principal Data Engineer & Critical Incident Manager



  • Lead the technical response to P0/P1 production incidents, coordinating investigation and resolution across teams.

  • Make time-critical decisions (rollback vs. forward-fix) and drive incidents to resolution within SLA.

  • Own root-cause analysis and postmortems, and assign and track follow-up action items.

  • Build and maintain monitoring, alerting, and on-call processes to detect and prevent incidents.

  • Author and improve incident runbooks, escalation paths, and severity classification.

  • Provide US-hours coverage for rapid response and mentor junior incident engineers.



  • Bachelor's degree in Computer Science, Engineering, or related field.

  • 6-9 years in SRE, production support, or platform engineering, including incident management.

  • Strong SQL and Python; experience with monitoring/alerting tools and cloud platforms (AWS/Azure).

  • Proven experience leading production incident response in a SaaS or multi-tenant environment.


LiftLab is the full-funnel MMM and incrementality testing platform that turns every dollar of brand and performance spend into compounding economic value. The platform combines Agile MMM, incrementality testing, and AI-powered scenario planning to help enterprise brands make better budget decisions with greater confidence and speed.


At the core of LiftLab is a closed-loop system: the Trust Engine continuously calibrates models through real-world geo experiments, PlatformSense catches daily ad platform shifts before they distort results, and the Scenario Planner translates model outputs into forward-looking budget decisions. The result is a platform that not only reports on past performance but also actively guides where the next marketing dollar should go.


LiftLab is trusted by marketing and finance leaders at brands including Pandora, SKIMS, Birkenstock, Cinemark, Anthropic, Hyundai, and Quicken. The company is SOC 2 compliant and ISO 27001 certified.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Incident Response Analyst (P0 Team)
Incident Response Analyst (P0 Team)

Liftlab, Inc. • Northern (KY)

Hybrid
USD 65,000 - 110,000
SRE Lead: Critical Incident Response (Remote)
SRE Lead: Critical Incident Response (Remote)

Liftlab, Inc. • Northern (KY)

Hybrid
USD 140,000 - 190,000
Remote Incident Response Analyst: P0 Troubleshooter
Remote Incident Response Analyst: P0 Troubleshooter

Liftlab, Inc. • Northern (KY)

Hybrid
USD 65,000 - 110,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
SRE
SRE

Ascendion • Jacksonville (FL), Northern (KY)

Hybrid
USD 125,000 - 136,000
Medical insurance
Dental insurance
Vision insurance
+7
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 190,000
Cloud Engineer
Cloud Engineer

TripleLift • New York (NY)

On-site
USD 130,000 - 170,000
Medical, Dental & Vision Plans
Flexible PTO
401k w/ employer match
+1
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000