Staff Site Reliability Engineer

Domino Data Lab

Argentina

Presencial

ARS 97.243.832 - 138.919.760

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Join Domino Data Lab as an SRE to lead the development of AI-assisted reliability tooling and improve incident response practices. You'll enhance observability coverage for critical systems while mentoring engineers and driving the reliability of our SaaS offerings. We look for deep experience in reliability engineering, fluency with Kubernetes and cloud platforms, and strong software engineering skills in Python or Go. This role fosters a culture of learning and growth, emphasizing diversity and collaboration.

Formación

  • Deep experience in Site Reliability Engineering or similar roles with hands‑on operational ownership.
  • Fluency with Kubernetes, Linux, and cloud platforms for production problem investigation.
  • Strong software engineering skills in Python or Go with a focus on building reliable tools.

Responsabilidades

  • Lead the development of AI‑assisted reliability tooling for incident resolution.
  • Own incident response end-to-end and improve documentation and understanding of problems.
  • Scale cloud operations practices for Domino’s SaaS offering.

Conocimientos

Site Reliability Engineering
Kubernetes
Linux
Cloud platforms
Observability tooling
Python
Go
Mentoring
Communication

Descripción del empleo

What We Are Building

As our infrastructure and customer footprint grow, we’re investing in a new kind of SRE practice where the people who respond to incidents also build the systems that make future incidents shorter, rarer, and less painful. We’re developing AI‑assisted tooling that helps our support and engineering teams diagnose problems faster, learn from outages more deeply, and automate away the toil that slows everyone down. This role sits at the center of that: equal parts hands‑on operator, software engineer, and technical leader. If you believe that operational experience and engineering craft make each other stronger, you’ll feel right at home here.

What Your Impact Will Be
  • Lead the development of Domino’s internal AI‑assisted reliability tooling, including systems that analyze tickets, logs, traces, and documentation to help teams resolve outages faster with less recurring toil
  • Improve the observability coverage and signal quality for our most critical customer‑facing systems, so engineers have more to work with throughout the development and support lifecycle
  • Own incident response end‑to‑end, from detection to remediation, and leave each problem space better documented, better understood, and less likely to recur
  • Guide the development of customer and user‑facing observability tools within our products
  • Define and mature SLO/SLI frameworks for priority services, turning abstract reliability goals into measurable, actionable standards
  • Scale cloud operations practices for Domino’s single‑tenant SaaS offering, and work with engineering teams to improve the reliability and repeatability of customer deployments and upgrades
  • Mentor other engineers and shape how SRE is practiced at Domino, including incident response workflows, operational readiness expectations, and post‑incident learning culture
What We Look For In This Role
  • Deep experience in Site Reliability Engineering, platform engineering, or a software engineering role with genuine, hands‑on operational ownership
  • Fluency with Kubernetes, Linux, cloud platforms, and observability tooling, and the ability to use them to investigate complex, real‑world production problems
  • A strong ability to perceive and close reliability gaps in technical products, tools and processes
  • Strong software engineering skills in Python or Go, with a track record of building internal tools or services that people actually rely on
  • Comfort leading technically ambiguous work and influencing direction across teams without needing direct authority to get things done
  • A history of improving reliability through engineering and automation, not just putting out fires manually
  • Strong communication skills and real experience mentoring engineers or shaping technical decision‑making on your team
  • Sound judgment about AI/LLM tooling: you know where it genuinely helps in operational workflows and where it adds noise instead of signal
  • Bonus: Experience with LLM‑based systems, retrieval workflows, SaaS platform operations, or building tooling for support or developer teams
What We Value
  • We strongly believe in the value of growing a diverse team and encourage people of all backgrounds, genders, ethnicities, abilities, and sexual orientations to apply
  • We value a growth mindset. High‑performing creative individuals who dig into problems and see the opportunities for success
  • We believe in individuals who seek truth and speak the truth and can be their whole selves at work
  • We value all of you that believe improving is always possible. At Domino, everything is a work in progress – we can do better at everything
  • We emphasize an environment of teaching and learning to equip employees with the tools needed to be successful in their function and the company
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Lead SRE for AI-Driven Reliability & Tooling
Lead SRE for AI-Driven Reliability & Tooling

Domino Data Lab • Argentina

Presencial
ARS 97.243.000 - 138.920.000
Senior Site Reliability Engineer/Platform Engineer
Senior Site Reliability Engineer/Platform Engineer

Techunting • Córdoba

Presencial
ARS 89.314.000 - 119.087.000
Senior Site Reliability Engineer IRC302878
Senior Site Reliability Engineer IRC302878

GlobalLogic • Municipio de Rincón de los Sauces

Presencial
ARS 178.643.000 - 223.304.000
Competitive salary
Family medical insurance
Extended paternity leave
+2
Senior Software Engineer, Patch Engineering & Vulnerability Remediation
Senior Software Engineer, Patch Engineering & Vulnerability Remediation

Domino Data Lab • Argentina

Presencial
ARS 119.095.000 - 208.417.000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

EPAM Systems • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+3
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

EPAM Systems • Argentina

Presencial
ARS 133.982.000 - 178.643.000
International projects
Global teams
LinkedIn Learning access
+2
Senior Site Reliability Engineer, AI-DNA, $100k/year USD
Senior Site Reliability Engineer, AI-DNA, $100k/year USD

IgniteTech • Argentina

Presencial
ARS 2.000.000 - 4.000.000
Site Reliability Engineer
Site Reliability Engineer

EPAM Systems • Argentina

Presencial
ARS 134.672.073 - 179.562.764
Lead Domino Developer - API
Lead Domino Developer - API

EPAM Systems • Argentina

Presencial
ARS 179.641.000 - 269.461.000
Senior Site Reliability Engineer IRC302878
Senior Site Reliability Engineer IRC302878

GlobalLogic • Buenos Aires

Híbrido
ARS 8.928.000 - 13.392.000
Exciting projects
Collaborative environment
Work-life balance
+2