Senior/Principal Application Support Engineer – AI Products & Agents

Recruise

Hyderabad

On-site

INR 3,500,000 - 7,000,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Recruise is seeking a Principal Application Support Engineer – Operations in Hyderabad to own complex incident resolution and drive systemic reliability across production services. You will mentor engineers, guide RCA efforts, and coordinate with engineering, product, QA, security, and platform teams to ensure resilient systems and strong operational readiness.

You will contribute to runbooks, monitoring, and automation to reduce toil, improve MTTR, and elevate customer experience.

Qualifications

  • 7–10 years in application support, production engineering, SRE, or software engineering with strong operations ownership and incident response.
  • Hands-on debugging across web apps, frontend/backend, integrations, and production environments.
  • Experience with incident management and ticketing workflows (ServiceNow, Jira) including major incident execution and RCA.
  • Strong knowledge of RESTful APIs and databases (e.g., PostgreSQL).
  • Experience with caching/data stores (Redis) and cloud platforms (AWS/Azure/GCP).
  • Expertise in monitoring/logging/alerting stacks (CloudWatch, ELK, Datadog, Splunk, AppDynamics).
  • Advanced scripting and automation (Bash, Python, JavaScript).
  • Experience in regulated industries and security-conscious operations.
  • Strong collaboration across Dev, QA, Product, Security, and Platform teams.
  • Proven mentoring and leadership to raise team capability.

Responsibilities

  • Act as final escalation point for complex, high-impact production issues across frontend, backend, integrations, data stores, and cloud infrastructure.
  • Lead major incident response with swarming, triage strategy, and cross-team coordination.
  • Drive RCA for recurring incidents and champion durable fixes over workarounds.
  • Translate RCA outcomes into durable changes in code, config, architecture, monitoring, or processes.
  • Define and implement runbooks, readiness checks, and rollback patterns with shift-left readiness.
  • Develop observability through logs, metrics, and traces; improve alert quality and dashboards.
  • Build automation for triage, remediation, and reporting to reduce toil.
  • Provide senior support for deployments including risk assessment and go/no-go input.
  • Collaborate with Dev/Platform teams to strengthen CI/CD safety and release checklists.
  • Mentor L2/R2 engineers and promote knowledge sharing across teams.

Skills

Incident management
Production debugging
SRE / production engineering
Automation scripting
Cross-team collaboration
Mentoring
Power Automate / Gen AI
Observability design
Cloud platforms
RESTful APIs
Security awareness
Python/Bash/JavaScript

Tools

ServiceNow
Jira
CloudWatch
ELK
Datadog
Splunk
AppDynamics
Power Automate
LLM tooling
PostgreSQL
Redis

Job description

Job Title: Principal Application Support Engineer – Operations

Industry: Healthcare / Pharmaceuticals

Seniority Level: Principal

Overview

Technology at Our Client builds and maintains capabilities using pioneering technologies like most prominent tech companies. What differentiates Technology at Our Client is that we create new possibilities through technology to advance our purpose – creating medicines that make life better for people around the world, like data-driven drug discovery and connected clinical trials. We hire technology professionals from a variety of backgrounds so they can bring an assortment of knowledge, skills, and diverse thinking to deliver solutions in every area of the business.

The Software Product Engineering (SPE) team is a specialized engineering group that delivers strategic solutions and differentiated capabilities. We take a forward-thinking approach, focusing on an enterprise platform and product mindset, ensuring that the solutions we build can be leveraged across Technology teams for broader impact and efficiency.

As a Principal Application Support Engineer – Operations, you will be the senior technical authority for production support across a suite of products and services. You will lead complex incident resolution, drive systemic reliability improvements, and influence operational standards across teams. This role expands beyond advanced troubleshooting to include end-to-end ownership of major incidents, deep technical remediation, automation to reduce operational toil, and mentoring of support engineers. You will partner closely with Engineering, Product, QA, Security, and Platform teams to ensure resilient services, strong operational readiness, and measurable improvements in uptime, latency, and customer experience.

Key Responsibilities
  • Act as the final escalation point for the most complex, high-impact production issues spanning frontend, backend, integrations, data stores, and cloud infrastructure.
  • Lead major incident response, including swarming/war-room execution, triage strategy, technical direction, and recovery coordination across multiple teams.
  • Drive consistent incident execution aligned with incident management expectations, including escalation, outage/deviation considerations, and appropriate stakeholder visibility.
  • Own and drive Root Cause Analysis (RCA) for recurring and severe incidents.
  • Identify systemic failure patterns and champion long‑term fixes over workarounds.
  • Partner with engineering to translate RCA outcomes into durable changes across code, configuration, architecture, monitoring, or process.
  • Track fixes to closure with measurable reliability impact.
  • Lead initiatives to improve availability, performance, scalability, and operational resilience, including reducing MTTR, improving detection, and reducing repeat incidents.
  • Define and implement operational guardrails, including readiness checks, runbooks, rollback patterns, post‑release validation, and shift‑left operational readiness with Dev/QE.
  • Contribute to or lead stabilization work consistent with engineering/SRE responsibilities, including reliability improvements, defect elimination, and major‑incident swarming.
  • Design and evolve observability across logs, metrics, and traces.
  • Improve signal quality through actionable alerts, noise reduction, and meaningful dashboards.
  • Build automation for common operational tasks, including triage, remediation, and reporting, using scripting and tooling to reduce manual effort and improve consistency.
  • Provide senior support for deployments and releases, including risk assessment, go/no‑go input, rollback readiness, and rapid response for post‑release issues.
  • Improve CI/CD operational safety through better validation, monitoring hooks, and release checklists in partnership with DevOps/Platform teams.
  • Ensure support processes and fixes align with internal standards and external regulations, including GDPR and HIPAA where applicable.
  • Promote secure operational practices, including least privilege, auditability, secure debugging, and appropriate handling of sensitive data during incident response.
  • Create and govern high‑quality runbooks, knowledge base articles, and operational standards.
  • Ensure reusability and adoption of operational knowledge and standards across teams.
  • Mentor L2/R2 engineers through technical coaching, incident handling patterns, RCA quality, and effective cross‑team collaboration.
  • Act as a role model for knowledge sharing.
Required Qualifications
  • 7–10 years of experience in application support, production engineering, SRE, or software engineering with strong operations ownership, including high‑severity incident response.
  • Deep hands‑on debugging experience across web applications, including frontend and backend, integrations, and production environments.
  • Strong experience with incident management and ticketing workflows, such as ServiceNow and Jira, including major incident execution and RCA.
  • Strong knowledge of RESTful APIs.
  • Strong knowledge of databases, such as PostgreSQL.
  • Strong knowledge of caching/data stores, such as Redis.
  • Strong knowledge of cloud platforms, including AWS, Azure, or GCP.
  • Expertise in monitoring, logging, and alerting stacks, such as CloudWatch, ELK, Datadog, Splunk, AppDynamics, or equivalent.
  • Ability to build actionable observability.
  • Advanced scripting and automation capability using technologies such as Bash, Python, or JavaScript to reduce toil and standardize response.
  • Experience supporting products in regulated industries.
  • Working knowledge of privacy and security expectations and secure handling.
  • Strong collaboration and communication skills across Dev, QA, Product, Security, and Platform teams.
  • Proven ability with skills such as Power Automate, LLM models, Gen AI, and Agentic Gen AI.
  • Demonstrated mentoring and leadership capability, including raising team capability through coaching and standards.
Preferred Qualifications
  • Experience defining and operationalizing SLIs/SLOs, error budgets, and reliability reporting using SRE ways of working.
  • Experience with containerization and deployment patterns, including Docker, Kubernetes, or ECS.
  • Experience with CI/CD systems.
  • Experience with infrastructure‑as‑code concepts.
What Success Looks Like
  • Be a recognized technical expert who solves complex problems and introduces improved methods and approaches for operations and reliability.
  • Lead technical decisions during incidents and influence operational standards, technical direction, and cross‑team alignment.
  • Demonstrate strong systems thinking and understand failure modes across distributed services, data stores, networks, and cloud infrastructure.
  • Drive measurable outcomes, including reduced repeat incidents, improved alert quality, lower MTTR, improved SLO attainment, and reduced manual toil.
  • Communicate crisply under pressure, facilitating fast alignment between engineering, product, and stakeholders during major incidents.
Additional Information
  • Availability to work flexible work hours may be required.
  • The team will support continuous operations across two shifts, and this role will require non‑standard work hours, including some work on weekends and holidays.
  • Appropriate adjustments in benefits will be provided for employees working non‑standard hours where applicable.
  • Candidates should be open to working different shifts when required:
    • 6:00 AM to 2:00 PM
    • 2:00 PM to 11:00 PM
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Optum • Hyderabad

On-site
INR 4,000,000 - 6,400,000
Application Development Associate Director (Production Support)
Application Development Associate Director (Production Support)

Evernorth Health Services • Hyderabad

On-site
INR 3,500,000 - 6,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Brillio • Bengaluru Urban

On-site
INR 1,200,000 - 2,000,000
Product Support Engineer
Product Support Engineer

Datavail Corp. • Mumbai

On-site
INR 2,800,000 - 6,000,000
Senior Application Support Operations Engineer
Senior Application Support Operations Engineer

Autonomize, Inc • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optum • Dadri

On-site
INR 2,500,000 - 4,500,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

UnitedHealth Group • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Lead Senior Production Support / Operations Engineer
Lead Senior Production Support / Operations Engineer

Datavail Career Site • Mumbai

On-site
INR 4,000,000 - 6,500,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Optum India • Hyderabad

On-site
INR 1,200,000 - 2,200,000
Senior Application Support Operations Engineer
Senior Application Support Operations Engineer

Autonomize AI • Bengaluru

On-site
INR 3,500,000 - 6,000,000