A complete application in a minute — tailored resume and cover letter, ready to send.
LiftLab, Inc. is seeking a Senior Site Reliability Engineer to join the Critical Incident Response team, leading rapid response to production-critical incidents across the multi-tenant platform.
The role collaborates with the Critical Incident Manager and Principal Data Engineer to drive incidents to resolution, implement postmortems, and strengthen platform reliability. Strong SQL/Python and cloud experience (AWS/Azure) are required, with US-hours coverage and mentorship responsibilities.
Senior engineer on the P0 / critical incident team, responsible for leading rapid response to production-critical (P0/P1) incidents across LiftLab's multi-tenant platform. Works directly with the Critical Incident Manager to drive incidents to resolution, minimize client impact, and strengthen the reliability of the platform.
6-9 yrs exp USA Remote Full-time Reports to: Principal Data Engineer & Critical Incident Manager
LiftLab is the full-funnel MMM and incrementality testing platform that turns every dollar of brand and performance spend into compounding economic value. The platform combines Agile MMM, incrementality testing, and AI-powered scenario planning to help enterprise brands make better budget decisions with greater confidence and speed.
At the core of LiftLab is a closed-loop system: the Trust Engine continuously calibrates models through real-world geo experiments, PlatformSense catches daily ad platform shifts before they distort results, and the Scenario Planner translates model outputs into forward-looking budget decisions. The result is a platform that not only reports on past performance but also actively guides where the next marketing dollar should go.
LiftLab is trusted by marketing and finance leaders at brands including Pandora, SKIMS, Birkenstock, Cinemark, Anthropic, Hyundai, and Quicken. The company is SOC 2 compliant and ISO 27001 certified.