Staff Site Reliability Engineer

Domino Data Lab

United States

On-site

USD 200,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Domino Data Lab is seeking a Site Reliability Engineer to enhance their internal AI-assisted reliability tooling. In this role, you will lead efforts to improve observability for critical systems and own incident response processes. Ideal candidates will have deep experience in SRE and strong software engineering skills, particularly in Python or Go.

The compensation range for this position is $200,000 - $230,000 USD. Join us to make a significant impact on our platform's reliability and performance!

Qualifications

  • Deep experience in Site Reliability Engineering or software engineering with operational ownership.
  • Fluency with Kubernetes, Linux, and cloud platforms for investigating production problems.
  • Strong software engineering skills in Python or Go.

Responsibilities

  • Lead the development of AI-assisted reliability tooling.
  • Own incident response from detection to remediation.
  • Guide customer-facing observability tools development.

Skills

Site Reliability Engineering
Kubernetes
Python
Go
Cloud platforms
Observability tooling
Communication skills

Job description

What Your Impact Will Be
  • Lead the development of Domino's internal AI-assisted reliability tooling, including systems that analyze tickets, logs, traces, and documentation to help teams resolve outages faster with less recurring toil
  • Improve the observability coverage and signal quality for our most critical customer-facing systems, so engineers have more to work with throughout the development and support lifecycle
  • Own incident response end-to-end, from detection to remediation, and leave each problem space better documented, better understood, and less likely to recur
  • Guide the development of customer and user-facing observability tools within our products
  • Define and mature SLO/SLI frameworks for priority services, turning abstract reliability goals into measurable, actionable standards
  • Scale cloud operations practices for Domino's single-tenant SaaS offering, and work with engineering teams to improve the reliability and repeatability of customer deployments and upgrades
  • Mentor other engineers and shape how SRE is practiced at Domino, including incident response workflows, operational readiness expectations, and post-incident learning culture
What We Look For In This Role
  • Deep experience in Site Reliability Engineering, platform engineering, or a software engineering role with genuine, hands-on operational ownership
  • Fluency with Kubernetes, Linux, cloud platforms, and observability tooling, and the ability to use them to investigate complex, real-world production problems
  • A strong ability to perceive and close reliability gaps in technical products, tools and processes
  • Strong software engineering skills in Python or Go, with a track record of building internal tools or services that people actually rely on
  • Comfort leading technically ambiguous work and influencing direction across teams without needing direct authority to get things done
  • A history of improving reliability through engineering and automation, not just putting out fires manually
  • Strong communication skills and real experience mentoring engineers or shaping technical decision-making on your team
  • Sound judgment about AI/LLM tooling: you know where it genuinely helps in operational workflows and where it adds noise instead of signal
  • Bonus: Experience with LLM-based systems, retrieval workflows, SaaS platform operations, or building tooling for support or developer teams

Compensation Range:
$200,000 - $230,000 USD

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE & Platform Engineer — AI-Driven Reliability
Senior SRE & Platform Engineer — AI-Driven Reliability

Socotra, Inc. • United States

On-site
USD 200,000 - 230,000
Equity
401(k) plan
Medical, dental, and vision benefits
+1
Senior SRE: AI-Driven Reliability & Observability
Senior SRE: AI-Driven Reliability & Observability

Domino Data Lab • United States

On-site
USD 200,000 - 230,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Principal Site Reliability Engineer
Principal Site Reliability Engineer

ViziRecruiter,LLC. • Quincy (MA)

Hybrid
USD 146,000 - 221,000
SRE Lead — AI-Driven Reliability & Observability (Remote)
SRE Lead — AI-Driven Reliability & Observability (Remote)

Domino Data Lab • United States

On-site
USD 120,000 - 150,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ViziRecruiter,LLC. • Quincy (MA)

Hybrid
USD 125,000 - 188,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Supio • San Francisco (CA)

On-site
USD 170,000 - 220,000