Senior Software Engineer - Incident Insights & Readiness

United States Digital Space LLC

Paris

Sur place

EUR 121 622 - 173 746

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

RSUs
ESPP
Professional development
Mentor program
Inclusive culture
Inclusion Talks
Mental health benefits
Global benefits

Résumé du poste

Datadog is seeking a seasoned Site Reliability Engineer to strengthen incident insights and readiness. You will own on-call practices, define incident response workflows, and contribute to post-mortems. The role emphasizes learning, blameless reviews, and cross-team collaboration to improve system resilience.

You will mentor engineers, lead cross-functional initiatives, and coach teams in incident management best practices. This hybrid role offers growth and a focus on operational excellence.

Qualifications

  • 5+ years building software that solves real user problems.
  • Experience designing new features and reviewing code/technical designs (Go/Python/TypeScript).
  • Experience building or operating distributed systems with Kubernetes and complex failure modes.
  • Ability to own ambiguous technical problems from design to delivery.
  • Experience analyzing incidents and driving improvements from operational learnings.
  • Experience in on-call rotations and incident management leadership.

Responsabilités

  • Own and improve the on-call experience by establishing best practices and platforms for rotations and compensation.
  • Define incident response workflows and lead software design to streamline processes.
  • Contribute to post-mortems and identify ways to reduce friction and increase learning value.
  • Support incident reviews and help share learnings across the organization to improve resilience.
  • Provide technical leadership and day-to-day coaching to team members.
  • Train on-call staff in incident and post-mortem processes.
  • Lead cross-functional initiatives across engineering to drive reliability and operational excellence.

Connaissances

Go
Python
TypeScript
Kubernetes
Incident management
On-call experience
Leadership
Collaboration
English communication

Outils

Go
Python
Kubernetes

Description du poste

The Incident Insights & Readiness SRE team at the company fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are the company engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across the company to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower the company to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement.

At the company, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.

The Incident Insights & Readiness SRE team at the company fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are the company engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across the company to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower the company to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement.

At the company, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.

What You’ll Do:
  • Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation.
  • Define how we respond to incidents, lead the design and implementation of software to streamline the process, and collaborate with product teams to improve incident response across the company. Our aim is to fully support our incident responders in dealing with complexity.
  • Contribute to the post-mortem process for the company, collaborating with teams on writing them, and identifying opportunities to reduce friction and enhance learning value for the organization. Our team also runs a weekly postmortem reading group.
  • Support various teams in facilitating incident reviews that emphasize learning and blamelessness. Help them share their learnings across the organization to improve the resilience of our people.
  • Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving and operational excellence best practices.
  • Train our on-callers in incident and post-mortem processes, sharing expertise in incident management best practices. This involves both introducing newcomers to on-call responsibilities and refreshing the knowledge of existing engineers.
  • Lead cross-functional initiatives in engineering organizations across the company, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence.
Who You Are:
  • At least 5 years of experience building software that solves real user problems. Experience designing new features and collaborating on code and technical design reviews. We primarily develop in Go and Python, with a bit of TypeScript.
  • Experience building or operating distributed systems, with familiarity with Kubernetes and an understanding of complex failure modes.
  • Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long‑term engineering quality with pragmatic execution.
  • Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings.
  • Experience participating in on‑call rotations and improving incident response processes. Experience serving as an incident commander or incident coordinator is a plus.
  • Empathy, collaboration, and communication skills in English to cultivate strong relationships across various teams in the organization
  • Experience mentoring engineers, driving cross‑functional initiatives, and influencing technical direction without relying on organizational authority.
  • We welcome candidates from a variety of backgrounds, including software engineering, site reliability engineering, production engineering, infrastructure, and other roles focused on building reliable systems or improving incident response.

the company values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply.

Benefits and Growth:
  • New hire stock equity (RSUs) and employee stock purchase plan (ESPP)
  • Continuous professional development, product training, and career pathing
  • Intradepartmental mentor and buddy program for in‑house networking
  • An inclusive company culture, ability to join our Community Guilds (the company employee resource groups)
  • Access to Inclusion Talks, our internal panel discussions
  • Free, global mental health benefits for employees and dependents age 6+
  • Competitive global benefits


Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with the company.

#LI-Hybrid

About the company:

the company is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers. Trusted globally by Fortune 500 companies and high‑growth AI leaders, the company enables businesses to move faster with clarity and confidence. Learn more about #DatadogLife on Instagram, LinkedIn, and the company Learning Center.

Equal Opportunity at the company:

the company is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and other characteristics protected by law. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. Here are our Candidate Legal Notices for your reference.

the company endeavors to make our Careers Page accessible to all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please complete

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior Software Engineer - Incident Insights & Readiness
Senior Software Engineer - Incident Insights & Readiness

Datadog • Paris

Hybride
EUR 140 000 - 190 000
Hybrid work model
Manager I, Engineering - Sensitive Data Scanner
Manager I, Engineering - Sensitive Data Scanner

United States Digital Space LLC • Paris

Hybride
EUR 110 000 - 150 000
Stock equity (RSUs)
Employee stock purchase plan (ESPP)
Professional development
+3
Manager I, Engineering - Code Security
Manager I, Engineering - Code Security

United States Digital Space LLC • Paris

Hybride
EUR 120 000 - 180 000
New hire stock equity (RSUs)
Employee stock purchase plan (ESPP)
Continuous professional development
+3
Manager I, Engineering - Source Code Integration
Manager I, Engineering - Source Code Integration

Datadog • Paris

Hybride
EUR 120 000 - 180 000
RSUs
ESPP
Career development
+3
Senior Software Engineer - Backend
Senior Software Engineer - Backend

Datadog • Paris

Hybride
EUR 50 000 - 75 000
New hire stock equity (RSUs)
Continuous professional development
Inclusive company culture
+1
Manager I, Engineering - Applied AI/ML Product Analytics Suite
Manager I, Engineering - Applied AI/ML Product Analytics Suite

United States Digital Space LLC • Paris

Hybride
EUR 120 000 - 160 000
Stock equity
Career development
Mentor programs
+1
Staff Engineer - Data Platform Experience
Staff Engineer - Data Platform Experience

Datadog • Paris

Hybride
EUR 130 000 - 210 000
Hybrid work arrangement
Employee stock purchase plan
RSUs
+3
Manager I, Engineering - Applied AI/ML Product Analytics Suite
Manager I, Engineering - Applied AI/ML Product Analytics Suite

Datadog • Paris

Hybride
EUR 150 000 - 190 000
RSUs
ESPP
Professional development
+4
Manager I, Engineering - Observability Pipelines (OP)
Manager I, Engineering - Observability Pipelines (OP)

Datadog • Paris

Hybride
EUR 120 000 - 180 000
New hire stock RSUs
Employee stock purchase plan (ESPP)
Professional development
+5
Manager I, Engineering - Security Libraries
Manager I, Engineering - Security Libraries

United States Digital Space LLC • Paris

Hybride
EUR 147 000 - 216 000
New hire stock equity (RSUs)
Employee stock purchase plan (ESPP)
Professional development & career path
+4