Incident Operations Lead

Jobgether

Lavamünd

Vor Ort

EUR 180.000 - 222.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine zielgenaue Bewerbung für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Competitive salary
Stock options
Health benefits
USD 500 home-office setup allowance
USD 150 monthly stipend

Zusammenfassung

Jobgether, partnering with a Switzerland-based client, seeks an Incident Operations Lead to build and helm a global incident-management function. You will define severity frameworks, run follow-the-sun coverage, and drive AI-powered automation across teams, ensuring fast, safe, and consistent incident responses.

You will lead a distributed team of Incident Commanders, shape governance, and collaborate with Risk, engineering, and communications to translate incident learnings into tangible

Qualifikationen

  • 5+ years in production engineering, SRE, or a closely related technical discipline.
  • Hands-on experience commanding high-severity incidents.
  • Experience building incident command or major-incident management function.
  • Experience designing severity models, escalation frameworks, and incident-response processes.
  • Ability to lead distributed teams across multiple time zones and operate a 24x7 on-call rotation.
  • Strong influence and stakeholder-management skills to drive action from diverse teams.
  • Understanding of reliability metrics and differentiating between productivity gains and real improvements.

Aufgaben

  • Lead and scale the incident operations function for the organization’s most critical incidents.
  • Establish 24x7 follow-the-sun coverage and escalation paths across regions.
  • Develop game days, tabletop exercises, simulations, and certification programs.
  • Build and implement AI-powered automation to reduce toil and speed incident response.
  • Coordinate with engineering, Risk, and customer communications teams for clear updates.
  • Define and track incident KPIs and post-incident actions with accountable owners.
  • Maintain the service catalog and ownership data to enable fast incident resolution.

Kenntnisse

Incident management
Team leadership
SRE
AI automation
Cross-timezone coordination

Jobbeschreibung

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an Incident Operations Lead based in Switzerland.

As an Incident Operations Lead, you will build and lead the function responsible for managing the most critical incidents across a complex, globally distributed technology environment. You will establish the operating model, severity framework, escalation paths, and 24x7 follow-the-sun coverage needed to respond effectively under pressure. The role sits at the intersection of engineering, SRE, Risk, communications, and reliability teams, ensuring every incident is managed with clarity and discipline. You will lead a distributed team of Incident Commanders while continuously improving response processes, metrics, and operational readiness. A major part of the role will be turning incident learnings into measurable reliability improvements and scalable operating practices. You will also shape the use of AI-powered automation to reduce operational toil and make incident management faster, safer, and more consistent.


Accountabilities
  • Build and lead a dedicated Incident Operations function responsible for coordinating the organization’s most critical incidents.
  • Recruit, develop, and certify Incident Commanders and establish a 24x7 follow-the-sun rotation across APAC, EMEA, and AMER.
  • Maintain effective regional handoffs and participate in the incident command roster to remain closely connected to operational realities.
  • Develop and run game days, tabletop exercises, simulations, and certification programs to keep responders prepared for high-severity events.
  • Establish a blameless incident culture focused on identifying process and system improvements rather than individual fault.
  • Own the incident severity model in partnership with Risk, ensuring severity levels appropriately reflect financial and regulatory materiality.
  • Define and maintain escalation paths, response thresholds, leadership escalation criteria, and procedures for unanswered alerts or pages.
  • Keep the service catalogue and service ownership information accurate and actionable so teams can quickly identify the right technical owners during incidents.
  • Coordinate the operational bridge between engineering and technical support teams resolving incidents and partner-facing communications teams providing customer updates.
  • Establish clear communication channels, provide accurate incident facts, and maintain appropriate update cadences throughout major incidents.
  • Lead post-incident retrospectives in collaboration with SRE and ensure follow-up actions are documented, assigned, prioritized, and delivered within defined service levels.
  • Track overdue post-incident reviews and actions by team and incident, ensuring every outstanding item has a clear next step and owner.
  • Define and maintain trusted operational KPIs, including time to respond and time to mitigate, with clear definitions and reliable underlying data.
  • Establish defensible performance baselines before setting improvement targets and use severity-based analysis to drive continuous improvement.
  • Build incident-management processes as scalable, documented, versioned operating products that can support significant organizational growth.
  • Lead an AI and agentic automation roadmap covering incident setup, timeline generation, RCA drafting, action tracking, update reminders, and review follow-ups.
  • Establish clear boundaries for automated workflows, defining where AI agents can operate independently and where human Incident Commander judgment remains essential.
  • Translate incident learnings into concise, actionable guidance and distribute them across engineering teams so lessons from individual failures can drive broader organizational improvement.
Requirements:
  • 5+ years of experience in production engineering, Site Reliability Engineering, technical operations, or a closely related technical discipline.
  • Proven hands-on experience commanding high-severity or major production incidents.
  • Demonstrated experience building or standing up an incident command or major-incident management function, rather than only participating in an established process.
  • Experience designing severity models, escalation frameworks, incident-response processes, and operational governance.
  • Proven ability to lead distributed teams across multiple time zones and operate a 24x7 on-call or follow-the-sun rotation.
  • Strong influence and stakeholder-management skills, with the ability to drive action from engineers and teams outside your direct reporting line.
  • Confident judgment when making and defending severity or escalation decisions during high-pressure situations.
  • Strong understanding of reliability metrics and the ability to distinguish improvements in measured performance from genuine improvements in operational outcomes.
  • Excellent scope discipline and the ability to establish clear ownership boundaries while ensuring issues are routed without creating operational gaps.
  • Strong written and verbal communication skills, including the ability to maintain calm and clarity during incident bridges.
  • Ability to communicate effectively with senior and executive stakeholders during high-impact incidents without minimizing or overstating the situation.
  • Strong understanding of FinTech environments and the trust, reliability, and operational requirements associated with API-driven financial platforms.
  • Experience using AI or agentic automation to reduce operational toil and improve incident-management workflows.
  • Excellent organizational skills and a strong focus on documentation, process quality, and continuous improvement.
  • Formal incident command, ITIL, Major Incident Management, or crisis-management training is a plus.
  • Experience running certification programs, game days, simulations, or operational drills is highly valued.
  • Familiarity with modern incident-management and on-call platforms is advantageous.
  • Experience building and maintaining service catalogues or service ownership registries is a plus.
  • Experience partnering with program management or reliability teams to convert incident follow-ups into funded improvement initiatives is desirable.
  • Familiarity with regulatory incident-reporting obligations such as DORA, Reg SCI, FINRA, or equivalent frameworks is advantageous.
  • Experience in online securities trading, capital markets, or another regulated and market-hours-sensitive environment is a plus.
  • Experience deploying an incident-management operating model across multiple regions or entities is advantageous.
Benefits:
  • Competitive salary.
  • Stock options.
  • Health benefits.
  • One‑time USD $500 home‑office setup allowance for new hires.
  • USD $150 monthly stipend provided through a company card.
  • Opportunity to work in a globally distributed environment.
  • Exposure to a highly technical, reliability‑focused operating environment.
  • Opportunity to lead and shape a critical operational function from the ground up.
  • Significant scope to influence incident management, reliability, automation, and organizational readiness.
  • Opportunity to work cross‑functionally with engineering, SRE, Risk, communications, and reliability teams.
  • Opportunity to develop and implement AI‑powered operational workflows and automation.
  • Inclusive environment that values curiosity, empathy, accountability, and diverse perspectives.

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre‑contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Global Incident Command Lead - 24/7 Reliability & AI Ops
Global Incident Command Lead - 24/7 Reliability & AI Ops

Jobgether • Lavamünd

Vor Ort
EUR 180.000 - 222.000
Competitive salary
Stock options
Health benefits
+2
Senior Backend Engineer, Core APIs (SRE Focus)
Senior Backend Engineer, Core APIs (SRE Focus)

Far Coder • Lavamünd

Hybrid
EUR 70.000 - 150.000
Fully remote
Globally distributed team
Mentorship and growth
DevOps Administrator
DevOps Administrator

Jobgether • Lavamünd

Vor Ort
EUR 95.000 - 132.000
Flexible work model
Remote onboarding
Performance bonuses
+5
Application Manager - Trading
Application Manager - Trading

Crypto Finance AG • Schweiz

Vor Ort
EUR 161.000 - 204.000
International team
Central Zurich workplace
Flat hierarchies
+1
Senior DevOps Engineer
Senior DevOps Engineer

Frontify • Lavamünd

Hybrid
EUR 127.000 - 191.000
25 days leave
Parental leave
Sickness benefits
+7
DevOps Administrator
DevOps Administrator

Jobgether • Österreich

Vor Ort
EUR 65.000 - 90.000
Flexible remote or office work
Remote onboarding
Performance-based bonuses
+5
SRE Engineer
SRE Engineer

QA Limited • Lavamünd

Hybrid
EUR 116.000 - 159.000
Paid vacation
Relocation assistance
Professional development
+3
Senior Solution Sales Executive (Risk & Security)
Senior Solution Sales Executive (Risk & Security)

ServiceNow • Wien

Vor Ort
EUR 90.000 - 130.000
Channel Solution Architect, DACH (Remote)
Channel Solution Architect, DACH (Remote)

CrowdStrike • Österreich

Vor Ort
EUR 90.000 - 130.000
Wellness programs
Parental leave
Career development
+2
Lead Software Developer (m/f/d)
Lead Software Developer (m/f/d)

Atos SE • Wien

Vor Ort
EUR 90.000