Principal Infrastructure Engineer - Major Incident Manager

svb

Bengaluru

On-site

INR 4,200,000 - 7,000,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

FC Global Services India LLP (First Citizens India) based in Bengaluru is hiring a Principal Infrastructure Engineer - Major Incident Manager to lead 24x7 incident response for critical enterprise applications, cloud platforms, and infrastructure.

The role requires acting as Incident Commander, driving technical triage, root‑cause analysis, and rapid recovery while partnering with Engineering, SRE, Cybersecurity, and Operations teams to improve platform reliability and operational resilience

Qualifications

  • Experience as Incident Commander for major incidents across enterprise applications, cloud platforms, infrastructure, databases, and network services.
  • Proven track record leading end-to-end major incident management activities including assessment, prioritization, escalation, engagement, coordination, recovery, and restoration.
  • Ability to drive technical triage and root-cause investigation under pressure, removing blockers and guiding recovery.
  • Familiarity with ITSM, operational risk, and regulatory requirements.

Responsibilities

  • Serve as the Incident Commander for major incidents and lead end-to-end response across platforms and services.
  • Direct technical bridge calls and coordinate among Infrastructure, Cloud, Network, Database, Middleware, and SRE teams.
  • Lead PIRs and RCA activities and drive corrective actions with Problem Management.
  • Develop and refine runbooks, playbooks, and escalation paths to reduce MTTR and improve resilience.
  • Produce executive incident updates and performance dashboards with key metrics.

Skills

Major Incident Management
Technical Leadership
Incident Commander
SRE Collaboration
Automation
Observability
Monitoring
Capacity Planning
Post-Incident Reviews

Job description

FC Global Services India LLP (First Citizens India), a part of First Citizens BancShares, Inc., a top 20 U.S. financial institution, is a global capability center (GCC) based in Bengaluru. Our India-based teams benefit from the company's over 125-year legacy of strength and stability. First Citizens India is responsible for delivering value and managing risks for our lines of business. We are particularly proud of our strong, relationship-driven culture and our long‑term approach, which are deeply ingrained in our talented workforce. This is evident across all key areas of our operations, including Technology, Enterprise Operations, Finance, Cybersecurity, Risk Management, and Credit Administration. We are seeking talented individuals to join us in our mission of providing solutions fit for our clients' greatest ambitions.

Job Description:
Value Proposition

Responsible for enhancing the reliability, resilience, and stability of enterprise IT services through effective Major Incident Management and rapid service restoration.

Leading the command, coordination, and technical response for critical production incidents impacting business-critical applications, infrastructure, cloud platforms, and customer-facing services.

Acting as the central Incident Commander during high‑severity incidents, driving technical triage, root‑cause identification, resolution, and minimizing business impact and operational risk.

Partnering with Engineering, SRE, Infrastructure, Cybersecurity, and Application teams to strengthen platform reliability, operational resilience, monitoring, automation, and service maturity.

Job Details

Position Title: Principal infrastructure Engineer - Major Incident Manager

Career Level: P4

Job Category: Assistant Vice President

Role Type: Hybrid

Job Location: Bangalore

About the Team:

The Major Incident Management Team serves as the central command function for critical incidents across global operations, providing 24x7 coverage through teams in India and the US.

Focused on rapid recovery and business continuity, the team leads coordinated incident response, promotes industry best practices, drives continuous improvement, and strengthens operational resilience to support a world‑class financial institution.

Key Deliverables (Duties and Responsibilities)
Major Incident Command & Technical Leadership

Serve as the Incident Commander for major incidents across enterprise applications, cloud platforms, infrastructure, databases, and network services.

Lead end‑to‑end major incident management activities including incident assessment, prioritization, escalation, responder engagement, technical coordination, recovery execution, and service restoration.

Direct technical bridge calls and facilitate collaboration among Infrastructure, Cloud, Network, Database, Middleware, Application Development, Cybersecurity, Vendor, and SRE teams.

Drive structured technical triage and ensure investigation efforts remain focused on business recovery and root‑cause isolation.

Challenge incomplete technical updates, validate remediation approaches, and remove operational blockers during critical situations.

Make informed decisions under pressure while maintaining clear ownership, accountability, and resolution momentum.

Ensure incidents are managed in accordance with established ITSM, operational risk, and regulatory requirements.

Service Reliability & SRE Partnership

Partner with Site Reliability Engineering (SRE), Production Support, and Engineering teams to improve service reliability and operational resilience.

Promote adoption of reliability engineering practices including:

  • Service Level Indicators (SLIs)
  • Service Level Objectives (SLOs)
  • Service Level Agreements (SLAs)
  • Error Budgets
  • Observability
  • Proactive Monitoring
  • Event Correlation
  • Capacity Planning
  • Operational Readiness Reviews

Identify recurring incidents, systemic weaknesses, and reliability risks across technology platforms.

Contribute to the continuous improvement of operational stability through automation, monitoring enhancements, and process optimization.

Support operational readiness for major releases, infrastructure transformations, and cloud migrations.

Incident Analysis & Continuous Improvement

Lead Post‑Incident Reviews (PIRs) and Root Cause Analysis (RCA) activities.

Partner with Problem Management teams to identify corrective and preventive actions.

Ensure action items are tracked through closure and measurable service improvements are achieved.

Drive improvements to runbooks, knowledge articles, operational procedures, escalation paths, and response frameworks.

Identify opportunities to reduce Mean Time to Detect (MTTD), Mean Time to Restore (MTTR), and recurring service disruptions.

Reporting & Operational Insights

Produce executive‑level incident communications and reporting.

Develop and publish operational dashboards and metrics including:

  • MTTD
  • MTTA
  • MTTR
  • Incident Volumes
  • Availability Metrics
  • Escalation Trends
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Infrastructure Engineer - Major Incident Manager
Principal Infrastructure Engineer - Major Incident Manager

Silicon Valley Bank • Karnataka

Hybrid
INR 2,500,000 - 4,500,000
Principal Infrastructure Engineer - Major Incident Manager
Principal Infrastructure Engineer - Major Incident Manager

First Citizens India • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Associate Director - SRE Support Engineering
Associate Director - SRE Support Engineering

svb • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Incident & Problem Manager
Incident & Problem Manager

Jobgether • India

Hybrid
INR 1,200,000 - 2,000,000
Flexible work arrangements
Professional development opportunities
Work-life balance emphasis
Associate Director - SRE Support Engineering
Associate Director - SRE Support Engineering

First Citizens India • Bengaluru

Hybrid
INR 4,000,000 - 7,500,000
Delivery Operation representative
Delivery Operation representative

Accenture in India • Gurugram District

On-site
INR 2,500,000 - 4,200,000
Major Incident Manager (Escalation Management Team)
Major Incident Manager (Escalation Management Team)

Genpact • Hyderabad

Hybrid
INR 1,800,000 - 3,000,000
Associate Director - SRE Support Engineering
Associate Director - SRE Support Engineering

Silicon Valley Bank • Karnataka

Hybrid
INR 4,000,000 - 7,000,000
Vice President, Enterprise Technology Command Center, Major Incident Manager, DTI Site Reliabil[...]
Vice President, Enterprise Technology Command Center, Major Incident Manager, DTI Site Reliabil[...]

DBS Bank • Hyderabad

On-site
INR 5,500,000 - 9,000,000
Principal Associate - IT (IT Incident Manager)
Principal Associate - IT (IT Incident Manager)

Eurofins • Coimbatore District

On-site
INR 1,400,000 - 2,200,000