Incident Response Engineer ||

Sprinklr

Bengaluru

On-site

INR 700,000 - 1,200,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sprinklr is seeking an experienced Incident Manager to own end-to-end production incident handling across a wide set of systems and teams in a fast-paced environment. You will triage, escalate, coordinate, and drive RCAs to prevent recurrence, while communicating clearly to technical and non-technical stakeholders.

The role requires strong governance with ITIL/ITSM practices, 24x7 readiness, and collaboration with Engineering, SRE, DevOps, and Infrastructure teams.

Qualifications

  • Experience in Incident Management, Major Incident Management, Production Support, or IT Service Management.
  • Hands-on experience managing P1/P2, Sev0/Sev1 incidents.
  • Strong understanding of ITIL and ITSM processes.
  • Experience working in SaaS, Cloud, Enterprise Software, or large-scale production environments.
  • Familiarity with monitoring and observability tools such as Grafana, Prometheus, Kibana/Elasticsearch, Splunk, Dynatrace, AppDynamics, or similar platforms.

Responsibilities

  • Take end-to-end ownership of production incidents from detection, logging, triage, impact assessment, prioritization, escalation, investigation, mitigation, resolution, recovery, and closure.
  • Act as the primary point of contact and single point of accountability for assigned incidents.
  • Lead P1/P2 and Sev0/Sev1 major incidents, establish incident bridges/war rooms, and ensure structured coordination across Engineering, SRE, DevOps, Infrastructure, Cloud, Database, Network, Product, Support, and other relevant teams.
  • Perform incident triage and determine the appropriate severity based on customer, business, and platform impact.
  • Coordinate with technical teams to identify the root cause, contributing factors, mitigation, and permanent corrective actions.
  • Provide timely, accurate, and structured communication to technical teams, business stakeholders, leadership, and customers as required.
  • Drive Root Cause Analysis (RCA) and Post-Incident Reviews (PIRs) for major incidents and ensure corrective and preventive actions are assigned, tracked, and closed.
  • Monitor production systems, service health dashboards, alerts, and operational tools to proactively identify customer-impacting issues.
  • Follow ITIL/ITSM-aligned Incident, Major Incident, Problem, and Change Management processes.
  • Coordinate with Change and Release Management teams to identify and minimize production risks associated with releases, maintenance, and infrastructure changes.
  • Develop, maintain, and continuously improve incident management SOPs, runbooks, escalation matrices, response procedures, communication templates, and operational documentation.
  • Track operational metrics including MTTD, MTTA, MTTR, response time, escalation adherence, incident SLA, RCA SLA, customer impact, and repeat incident trends.
  • Prepare incident reports, operational dashboards, management summaries, and trend analysis to drive continuous improvement.

Skills

Incident management
Major incident handling
SRE coordination
Problem management
Stakeholder communication

Tools

Jira
Confluence
PagerDuty
Grafana
Prometheus
Kibana/Elasticsearch
Azure DevOps
GitHub

Job description

Sprinklr is the definitive, AI-native platform for Unified Customer Experience Management (Unified-CXM), empowering brands to deliver extraordinary experiences at scale — across every customer touchpoint.

By combining human instinct with the speed and efficiency of AI, Sprinklr helps brands earn trust and loyalty through personalized, seamless, and efficient customer interactions. Sprinklr’s unified platform provides powerful solutions for every customer-facing team — spanning social media management, marketing, advertising, customer feedback, and omnichannel contact center management — enabling enterprises to unify data, break down silos, and act on real-time insights.

Today, 1,900+ enterprises and 60% of the Fortune 100 rely on Sprinklr to help them deliver consistent, trusted customer experiences worldwide.

Job Description
  • Take end-to-end ownership of production incidents from detection, logging, triage, impact assessment, prioritization, escalation, investigation, mitigation, resolution, recovery, and closure.
  • Act as the primary point of contact and single point of accountability for assigned incidents.
  • Lead P1/P2 and Sev0/Sev1 major incidents, establish incident bridges/war rooms, and ensure structured coordination across Engineering, SRE, DevOps, Infrastructure, Cloud, Database, Network, Product, Support, and other relevant teams.
  • Perform incident triage and determine the appropriate severity based on customer, business, and platform impact.
  • Ensure the right technical owners and subject-matter experts are engaged through the defined escalation matrix.
  • Drive incidents with clear ownership, timelines, action items, escalation paths, and next steps until service restoration.
  • Analyze application failures, service degradation, infrastructure issues, API failures, database issues, distributed-system problems, cloud incidents, and other production issues.
  • Review application logs, monitoring dashboards, alerts, service metrics, and system dependencies to support incident investigation.
  • Develop a good understanding of application architecture, APIs, microservices, databases, distributed systems, cloud infrastructure, and service dependencies to effectively drive technical discussions.
  • Coordinate with technical teams to identify the root cause, contributing factors, mitigation, and permanent corrective actions.
  • Provide timely, accurate, and structured communication to technical teams, business stakeholders, leadership, and customers as required.
  • Communicate incident impact, investigation progress, mitigation status, recovery updates, and next steps clearly to both technical and non-technical stakeholders.
  • Ensure communication follows defined templates, frequency, SLA, and escalation standards.
  • Drive Root Cause Analysis (RCA) and Post-Incident Reviews (PIRs) for major incidents and ensure corrective and preventive actions are assigned, tracked, and closed.
  • Analyze recurring incidents and work with Problem Management and Engineering teams to eliminate repeat issues and improve service reliability.
  • Monitor production systems, service health dashboards, alerts, and operational tools to proactively identify customer-impacting issues.
  • Work with Monitoring, SRE, and Engineering teams to improve alert quality, thresholds, detection coverage, dashboards, and early-warning mechanisms.
  • Identify monitoring gaps and recommend new alerts, dashboards, automation, correlation, and auto-remediation opportunities.
  • Review incident and alert trends to identify opportunities for alert reduction, tuning, proactive detection, and operational automation.
  • Follow ITIL/ITSM-aligned Incident, Major Incident, Problem, and Change Management processes.
  • Coordinate with Change and Release Management teams to identify and minimize production risks associated with releases, maintenance, and infrastructure changes.
  • Support service transition activities and ensure operational readiness before services are moved into production.
  • Develop, maintain, and continuously improve incident management SOPs, runbooks, escalation matrices, response procedures, communication templates, and operational documentation.
  • Track operational metrics including MTTD, MTTA, MTTR, response time, escalation adherence, incident SLA, RCA SLA, customer impact, and repeat incident trends.
  • Prepare incident reports, operational dashboards, management summaries, and trend analysis to drive continuous improvement.
  • Ensure effective shift handovers and continuity of ongoing incidents in a 24×7 global operations model.
  • Coordinate resources during critical incidents, assign tasks, track action items, and ensure the appropriate teams are engaged at the right time.
  • Support junior Incident Managers and team members by providing guidance on incident handling, escalation, communication, and operational processes.
  • Coordinate with external vendors and technology partners during incidents involving third-party services and ensure timely updates, mitigation, RCA, and accountability.
  • Identify opportunities to improve service reliability, incident response, monitoring, automation, operational efficiency, and customer experience.
  • Requirements
  • 3+ years of experience in Incident Management, Major Incident Management, Production Support, Technical Operations, NOC, SRE, Application Support, or IT Service Management.
  • Hands-on experience managing critical production incidents such as P1/P2, Sev0/Sev1, or equivalent severity incidents.
  • Strong understanding of Incident Management, Problem Management, Change Management, ITIL, and ITSM processes.
  • Experience working in SaaS, Cloud, Enterprise Software, or large-scale production environments.
  • Good understanding of application architecture, service dependencies, APIs, microservices, databases, distributed systems, and cloud technologies.
  • Working knowledge of AWS, Azure, or GCP.
  • Strong troubleshooting, analytical, problem-solving, and decision-making skills.
  • Excellent verbal and written communication skills with the ability to manage technical, business, leadership, and customer stakeholders.
  • Strong escalation, coordination, and stakeholder management skills.
  • Ability to remain calm, structured, and effective during high-pressure critical incidents.
  • Ability to manage multiple incidents, priorities, and stakeholders simultaneously.
  • Experience working in a 24×7 operations environment.
  • Experience with monitoring and observability tools such as Grafana, Prometheus, Kibana, Elasticsearch, Splunk, Dynatrace, AppDynamics, or similar platforms.
  • Experience with ITSM and incident management tools such as Jira, Confluence, PagerDuty, or equivalent platforms.
  • Familiarity with Azure DevOps and GitHub.
  • Awareness of automation, AIOps, intelligent alerting, incident correlation, and auto-remediation concepts.
  • Experience with customer communication, operational reporting, and vendor coordination is preferred.
  • ITIL Foundation certification or equivalent knowledge is preferred.
  • Strong proficiency in MS Excel, PowerPoint, and Visio for reporting, dashboards, presentations, and process documentation.
We focus on our mission

Sprinklr was founded in 2009 to solve a big problem: growing enterprise complexity that separated brands from their customers. Our vision was clear: to unify fragmented teams, tools and data — helping large organizations build deeper, more meaningful connections with the people they serve.

Today, Sprinklr has a unified, AI-native platform for four product suites: Sprinklr Service, Sprinklr Social, Sprinklr Marketing, and Sprinklr Insights.

Sprinklr is here to do three things:

  • Lead a new category of enterprise software that we call Unified-CXM.
  • Empower companies to deliver next generation, unified engagement journeys that reimagine the customer experience.
  • Create a culture of customer obsession, with trust, teamwork, and accountability.
We believe in our product

Customers who value exceptional customer experiences have what they need on our single unified platform, built with an operating system approach on a single codebase. That means that everything — and everyone — can work together to service, respond, sell, and market to customers on the channels they prefer. While Unified Customer Experience Management (Unified-CXM) as a category is just getting started, we are well on our way to creating a no-compromise, unified approach to better customer experiences for the world’s leading enterprise brands.

We invest in our people

We offer a comprehensive suite of benefits designed to help each member of our team thrive. Sprinklr believes that you should be able to get the type of care you need for your personal well-being when you need it. We believe it is important to take time off – it is essential for your mental and physical wellbeing. We provide Sprinklrites with paid time off to recharge and spend time with loved ones. We want to grow our talent with purpose. Our open Mentoring Program is designed to create meaningful connections that support growth and amplify our focus.

EEO - Our philosophy

Our goal is to ensure every employee feels like they belong and are operating in a collaborative environment. We fervently believe every employee matters and should be respected and heard. We believe we are stronger when we belong because collectively, we’re more innovative, creative, and successful.

Sprinklr is proud to be an equal-opportunity workplace and complies with all applicable federal, state, and local fair employment practices laws. We are committed to equal employment opportunity regardless of race, color, religion, creed, national origin or ancestry, ethnicity, sex (including gender, pregnancy, sexual orientation, and gender identity), age, physical or mental disability, citizenship, past, current, or prospective service in the uniformed services, genetic information, or any other characteristic protected under applicable law.

Reasonable accommodations are available upon request during the interview process. To request an accommodation, please work directly with your recruitment coordinator or recruiter.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Incident Response Engineer |
Incident Response Engineer |

Sprinklr • Bengaluru

On-site
INR 1,200,000 - 2,600,000
Paid time off
Mentoring program
Strategic Technical Program Manager
Strategic Technical Program Manager

Sprinklr • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Sr. Customer Insights Analyst
Sr. Customer Insights Analyst

Sprinklr • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Support Platform Administrator
Support Platform Administrator

Sprinklr • Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
Paid time off
Mentoring program
Comprehensive benefits
Support Platform Administrator
Support Platform Administrator

Sprinklr • Gurugram District

On-site
INR 1,200,000 - 2,000,000
Sr. Managed Services Consultant (AMER Shift)
Sr. Managed Services Consultant (AMER Shift)

Sprinklr • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Sr. Business Analyst
Sr. Business Analyst

Sprinklr • Gurugram District

On-site
INR 900,000 - 1,300,000
Senior Technical Program Manager
Senior Technical Program Manager

Sprinklr • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Senior Product Security Engineer
Senior Product Security Engineer

Sprinklr India Private Limited • Gurugram District

On-site
INR 2,500,000 - 4,500,000
Lead Managed Services Consultant
Lead Managed Services Consultant

Sprinklr • Gurugram District

On-site
INR 900,000 - 1,500,000