Senior Software Engineer, Chaos Engineering

United States Digital Space LLC

Paris

Hybride

EUR 90 000 - 130 000

Plein temps

Il y a 42 heures
Soyez parmi les premiers à postuler
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV personnalisé et une lettre de motivation en environ une minute.

Passez les filtres ATS

Résumé du poste

Datadog is seeking a Senior Software Engineer to advance chaos engineering and zonal resilience across production systems. You will build automation to evacuate and recover from zonal failures while contributing to fault injection, incident replay, and reliability tooling.

Strong skills in distributed systems, Kubernetes, and gRPC are essential for collaborating with multiple teams. You will help advance AI-assisted operational workflows and play a key role in designing safe, automated

Qualifications

  • Strong distributed systems fundamentals with knowledge of failure modes and recovery.
  • Experience with Kubernetes workload lifecycles and resilient design.
  • Background in building production systems with safety and controlled failure handling.
  • Clear communication of complex technical decisions in runbooks and design docs.

Responsabilités

  • Build zonal-resilience automation for safe evacuations, switchovers, and recovery.
  • Design and build fault-injection systems for production environments and controlled experiments.
  • Develop safeguards such as blast-radius controls, kill switches, and rollback paths.
  • Create agents and automation to propose failure scenarios and track remediation.
  • Lead gamedays from hypothesis design through execution and findings.

Connaissances

Distributed systems
Kubernetes
gRPC
Cloud infrastructure
Observability

Outils

Kubernetes controllers
CI/CD pipelines
Incident replay tooling

Description du poste

the company's Chaos Engineering team builds systems that surface reliability weaknesses before they become outages. As a Senior Software Engineer, you will initially focus on zonal resilience, building automation that helps services safely evacuate and recover from zonal failures, while also contributing to fault injection, incident replay, gameday orchestration, and reliability tooling. You will work across engineering teams to design systems that safely exercise production failure modes and turn findings into verified remediation. You will also help advance the use of AI and automation to identify, test, and close resilience gaps as the company's software and infrastructure evolve.

At the company, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.

What You'll Do:
  • Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services.
  • Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments.
  • Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible.
  • Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification.
  • Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes.
  • Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers.
Who You Are:
  • You have strong distributed systems fundamentals and can reason about consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery.
  • You understand Kubernetes workload lifecycles, including how pods, controllers, scheduling, draining, and eviction interact with resilient system design.
  • You have experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important.
  • You communicate complex technical decisions clearly through design documents, runbooks, postmortems, and cross-functional technical discussions.
  • You are comfortable collaborating across engineering teams to understand unfamiliar systems, identify failure modes, and drive resilience improvements.
  • Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems is beneficial but not required.

the company values people from all walks of life. We know not everyone will meet all the above qualifications on day one. That's okay. If you're passionate about technology and want to grow your experience, we encourage you to apply.

Benefits and Growth:
  • Develop deep expertise in distributed systems, production resilience, Kubernetes, and large-scale infrastructure.
  • Work on reliability systems that operate across the company’s production environment and influence how engineering teams design for failure.
  • Grow your experience designing safe, automated approaches to fault injection, zonal resilience, and incident reproduction.
  • Explore practical applications of AI and automation to reliability engineering and operational workflows.
  • Collaborate with engineers across infrastructure, databases, observability, and service teams on complex systems challenges.
  • Mentor other engineers and contribute to technical designs, engineering practices, and platform strategy.
  • Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with the company.

#LI-Hybrid

About the company:

the company is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers. Trusted globally by Fortune 500 companies and high-growth AI leaders, the company enables businesses to move faster with clarity and confidence. Learn more about #DatadogLife on Instagram, LinkedIn, and the company Learning Center.

Equal Opportunity at the company:

the company is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and other characteristics protected by law. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. Here are our Candidate Legal Notices for your reference.

Privacy and AI Guidelines:

Any information you submit to the company as part of your application will be processed in accordance with the company'sApplicant and Candidate Privacy Notice. For information on our AI policy, please visit Interviewing at the company AI Guidelines.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior Software Engineer - Distributed Systems
Senior Software Engineer - Distributed Systems

Datadog • Bordeaux

Sur place
EUR 90 000 - 130 000
RSUs
ESPP
Career development
+3
Manager I, Engineering - Applied AI/ML Product Analytics Suite
Manager I, Engineering - Applied AI/ML Product Analytics Suite

Datadog • Paris

Sur place
EUR 150 000 - 190 000
RSUs
ESPP
Professional development
+4
Manager I, Engineering - Code Security
Manager I, Engineering - Code Security

Datadog • Paris

Sur place
EUR 120 000 - 180 000
New hire stock equity (RSUs)
Employee stock purchase plan (ESPP)
Continuous professional development
+4
Senior Software Engineer - Backend
Senior Software Engineer - Backend

Datadog, Inc. • Bordeaux

Sur place
EUR 90 000 - 130 000
RSUs and ESPP equity
Career development and product培训
Mentor and buddy programs
+1
Senior Software Engineer - Backend
Senior Software Engineer - Backend

Datadog • Paris

Sur place
EUR 50 000 - 75 000
New hire stock equity (RSUs)
Continuous professional development
Inclusive company culture
+1
Senior Application Security Engineer
Senior Application Security Engineer

Datadog, Inc. • Paris

Hybride
EUR 120 000 - 160 000
RSUs
ESPP
Professional development
+2
Senior Software Engineer - Distributed Systems
Senior Software Engineer - Distributed Systems

Datadog • Valbonne

Sur place
EUR 90 000 - 130 000
RSUs
ESPP
Career development
+2
Senior Software Engineer - Distributed Systems
Senior Software Engineer - Distributed Systems

Datadog • Nantes

Sur place
EUR 70 000 - 120 000
New hire stock equity (RSUs) and ESPP
Continuous professional development, 1
Intradepartmental mentor and buddy
+4
Senior Software Engineer - Distributed Systems
Senior Software Engineer - Distributed Systems

Datadog • Paris

Sur place
EUR 90 000 - 150 000
RSUs
ESPP
Professional development
+3
Senior Software Engineer - Incident Insights & Readiness
Senior Software Engineer - Incident Insights & Readiness

Datadog • Paris

Sur place
EUR 110 000 - 140 000
RSUs and ESPP
Career development
Mentor and buddy program
+2