Senior Incident Manager

lambda

Wintzenheim

Sur place

EUR 90 000 - 130 000

Plein temps

14 jours+
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV personnalisé et une lettre de motivation en environ une minute.

Passez les filtres ATS

Avantages offerts par ce poste

Health coverage
Wellness stipend
Commuter stipend

Résumé du poste

Lambda, The Superintelligence Cloud, is seeking a Senior Incident Manager to lead critical incident response across AI data center infrastructure. You will act as Incident Commander during major outages and coordinate cross-functional teams.

The role requires 8+ years in incident management, SRE, or infrastructure operations, with strong communication and experience with PagerDuty, ServiceNow, Jira, Datadog, and observability tooling. On-call rotation is included.

Qualifications

  • 8+ years of experience in incident management, site reliability engineering, or infrastructure operations.
  • Experience managing incidents in large-scale distributed infrastructure environments.
  • Strong understanding of data center operations, GPU compute clusters, networking and storage infrastructure, and cloud/hybrid platforms.

Responsabilités

  • Lead the response to critical incidents (SEV-1/SEV-2) impacting AI infrastructure and data center operations.
  • Act as Incident Commander during major outages coordinating multiple teams.
  • Provide updates to leadership and stakeholders during incidents and post-incidents.
  • Own the incident response lifecycle from triage to resolution and PIRs.
  • Maintain incident documentation, runbooks, and dashboards; drive continuous improvement.

Connaissances

Incident management
Site reliability
Infrastructure operations
Leadership
Communication
Stakeholder management

Outils

PagerDuty
ServiceNow
Jira
Datadog
Prometheus
Grafana

Description du poste

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

We are seeking a Senior Incident Manager to lead critical incident response across our AI data center infrastructure. This role is responsible for coordinating rapid resolution of service-impacting events, improving operational resilience, and driving incident management best practices across infrastructure, networking, platform engineering, and data center operations.

Role Overview

The Senior Incident Manager is responsible for leading the end-to-end lifecycle of operational incidents impacting AI infrastructure and data center services. This individual acts as the central command point during major incidents, ensuring rapid triage, cross-team coordination, effective communication, and structured post-incident analysis.

This role requires deep operational expertise in high-availability infrastructure, large-scale GPU clusters, networking, and cloud platforms, along with strong leadership and communication skills.

What You’ll Do
Incident Leadership
  • Lead the response to critical (SEV-1 / SEV-2) incidents impacting AI infrastructure, GPU clusters, networking, storage, and data center operations.

  • Serve as the Incident Commander during major outages, coordinating engineering, networking, facilities, and vendor teams.

  • Act as the liaison between leadership and external teams during incidents / post-incidents to provide updates and status summaries.

  • Establish clear incident timelines, triage actions, and resolution plans.

Incident Management Operations
  • Own the incident response lifecycle including:

    • Assisting Technical Triage

    • Escalation

    • Coordination

    • Resolution
      Post-incident review

  • Ensure timely and accurate communication with internal stakeholders and leadership.

  • Maintain incident response documentation and operational playbooks.

  • Conduct analysis on incidents and identify patterns / trends for improvement in response and systems reliability.

  • Work in an On-Call Rotation to respond to, lead, and coordinate incidents

Cross-Functional Coordination
  • Work closely with:

    • Data center operations

    • Infrastructure engineering & operations

    • Network engineering

    • Platform reliability engineering

    • Security operations

    • Hardware and facility vendors

  • Drive alignment during outages involving multiple infrastructure layers.

Post-Incident Analysis & Continuous Improvement
  • Lead post-incident reviews (PIRs) and root cause analysis. Identify systemic reliability gaps and implement corrective actions.

  • Track incident metrics including MTTR, MTTD, and incident recurrence rates.

Operational Excellence
  • Improve incident response processes, escalation paths, and tooling by working with technical support and engineering teams..

  • Contribute to runbooks, operational standards, and reliability frameworks.

  • Support implementation of automation and observability improvements.

Communication & Reporting
  • Provide executive-level incident summaries and reports.

  • Deliver clear, concise updates during active incidents.

  • Maintain incident dashboards and operational health reporting.

You
  • 8+ years experience in incident management, site reliability engineering, or infrastructure operations

  • Experience managing incidents in large-scale distributed infrastructure environments

  • Strong understanding of:

    • Data center operations

    • GPU compute clusters
      Networking and storage infrastructure

    • Cloud or hybrid infrastructure platforms

  • Proven ability to lead high-pressure incident response situations

  • Experience with incident management frameworks (ITIL, SRE, or equivalent)

  • Excellent communication and stakeholder management skills

  • Experience with incident tracking and monitoring tools such as:

    • PagerDuty

    • ServiceNow

    • Jira

    • Datadog

    • Prometheus / Grafana

Nice to Have
  • Experience operating AI or HPC infrastructure

  • Background in SRE, infrastructure engineering, or data center operations

  • Familiarity with high-density GPU environments (NVIDIA clusters, InfiniBand networks)

  • Experience with hyperscale or colocation data center environments

  • Knowledge of automation and incident response tooling

  • Knowledge of and experience with Incident command system (ICS)

  • Experience in leading and developing incident command from stractch

Key Competencies
  • Incident Command & Leadership

  • Operational Decision Making

  • Cross-Team Coordination

  • Root Cause Analysis

  • Crisis Communication

  • Infrastructure Reliability

What Success Looks Like in This Role
  • Reduced Mean Time to Resolution (MTTR) for critical incidents

  • Improved cross-team incident coordination

  • High-quality post-incident reviews and corrective actions

  • Increased infrastructure reliability and operational maturity

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda
  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Data Center Standards Architect
Data Center Standards Architect

lambda • Wintzenheim

Sur place
EUR 211 000 - 316 000
Health coverage
Dental coverage
Vision coverage
+4
Data Center Operations Systems Engineer (San Jose)
Data Center Operations Systems Engineer (San Jose)

lambda • Wintzenheim

Sur place
EUR 55 000 - 75 000
Health, dental, and vision coverage
Commuter stipend
401k plan with company match (USA)
+1
Senior HPC Platform Hardware Engineer
Senior HPC Platform Hardware Engineer

lambda • Wintzenheim

Sur place
EUR 158 000 - 228 000
Health, dental, and vision coverage
Wellness stipend
401k match (USA)
Staff Connectivity Engineer
Staff Connectivity Engineer

lambda • Wintzenheim

Sur place
EUR 132 000 - 167 000
Cash & equity compensation
Health, dental, and vision coverage
Wellness and commuter stipends
+2
Senior Incident Manager — AI Cloud Infra (Equity)
Senior Incident Manager — AI Cloud Infra (Equity)

lambda • Wintzenheim

Sur place
EUR 90 000 - 130 000
Health coverage
Wellness stipend
Commuter stipend
Data Center IT Technician
Data Center IT Technician

United States Digital Space LLC • Béthune

Sur place
EUR 38 000 - 48 000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • France

Sur place
EUR 90 000 - 130 000
Incident Operations Lead
Incident Operations Lead

Jobgether • France

Sur place
EUR 120 000 - 190 000
Stock options
Health benefits
One-time USD 500 home-office setup
+1
Compute Infrastructure Lead Onsite (Paris, France)
Compute Infrastructure Lead Onsite (Paris, France)

S27a • Paris

Hybride
EUR 120 000 - 160 000
DC IT Support Manager
DC IT Support Manager

Nebius • Béthune

Sur place
EUR 55 000 - 75 000
Competitive salary
Comprehensive benefits package
Flexible working arrangements