Lead Senior Production Support / Operations Engineer

Datavail Infotech

Mumbai

On-site

INR 1,800,000 - 3,000,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Datavail Infotech seeks a Lead Senior Production Support / Operations Engineer to own and enhance the in-house monitoring and observability platform across enterprise databases. You will serve as the operational bridge between product development, ITSM, and DBA teams, ensuring high reliability across Oracle, MySQL, PostgreSQL, MongoDB, Cassandra, and SQL Server environments.

You will lead Production Support Engineers, coordinate incident responses, and drive improvements in monitoring quality,

Qualifications

  • 8+ years of production support experience in enterprise environments.
  • Strong leadership experience in supervising Production Support Engineers.
  • Excellent cross-team collaboration and communication skills.

Responsibilities

  • Provide operational ownership for the in-house enterprise monitoring platform.
  • Monitor health, performance, alerts, and stability of monitoring services.
  • Coordinate incident triage across DBA, monitoring, infrastructure, and Service management teams.
  • Lead RCA and post-incident reviews.

Skills

Production leadership
Monitoring & observability
Cloud platforms
Windows & Linux
ServiceNow
Incident management
24x7 support

Education

Any Graduation

Tools

Prometheus
Grafana
Datadog

Job description

Job Title: Lead Senior Production Support / Operations Engineer
Education: Any Graduate
Experience: 8+years
Job Location: Mumbai
Role Summary

We are seeking a highly Senior Production Support / Operations Engineer to represent and operationally support our in-house monitoring and observability platform across enterprise database environments.

This role acts as the critical operational bridge between:

  • The in-house monitoring product development and QE team(s),
  • Internal ITSM/ticketing teams using ServiceNow,
  • Multiple enterprise DBA organizations across:
  • Oracle Database
  • Open-source database platforms (MySQL, PostgreSQL, MongoDB, Cassandra, etc.)
  • Microsoft SQL Server

The ideal candidate combines strong operational troubleshooting skills, production support leadership, monitoring expertise, incident coordination capabilities, and excellent cross-team collaboration skills in large-scale enterprise environments.

This is a senior leadership position: the candidate must bring substantial experience leading Production Support Engineers, combined with exceptional communication skills for direct, confident engagement with customers and senior stakeholders of in-house management.

Key Responsibilities
Production Operations & Monitoring Support
  • Provide operational ownership and production support for the in-house enterprise monitoring platform.
  • Monitor health, performance, alerting quality, and operational stability of monitoring services.
  • Analyze monitoring gaps, false positives, missed alerts, and operational inefficiencies.
  • Ensure monitoring coverage across Oracle, Open-source, and SQL Server database environments.
Incident & Escalation Management
  • Act as the operational point-of-contact during production incidents involving monitoring failures, alerting gaps, or infrastructure issues.
  • Coordinate incident triage across:
  • DBA teams,
  • Monitoring development teams,
  • Infrastructure teams,
  • Service management teams.
  • Drive bridge calls and ensure effective stakeholder communication during critical outages.
  • Perform root cause analysis (RCA) and post-incident operational reviews.
ServiceNow & Ticket Workflow Coordination
  • Work with ServiceNow for:
  • Operational escalations,
  • Service requests.
  • Review ticket quality and ensure operational accuracy of issue classification and routing.
  • Improve ticket workflows between DBA teams and monitoring platform support teams.
  • Collaborate with internal support organizations to streamline escalation processes.
Cross-Functional DBA Collaboration
  • Collaborate closely with enterprise DBA teams supporting:
  • Oracle Database
  • MySQL
  • PostgreSQL
  • MongoDB
  • Apache Cassandra
  • Microsoft SQL Server
  • Cloud services (AWS, AZURE, GCP)
  • Understand operational monitoring requirements specific to each database technology.
  • Work with DBAs to validate alert thresholds, event correlation, and monitoring accuracy.
  • Serve as the operational liaison between DBAs and monitoring team developers and QE.
Operational Excellence & Reliability Engineering
  • Identify recurring operational pain points and recommend automation opportunities.
  • Improve alert quality, event correlation, and monitoring reliability.
  • Participate in operational readiness reviews for new monitoring features.
  • Help define operational standards, playbooks, and escalation procedures.
Monitoring & Observability Engineering
  • Support enterprise observability initiatives involving:
  • Metrics,
  • Events,
  • Alerting,
  • Dashboards,
  • Health monitoring,
  • Incident correlation.
  • Work with both commercial and in-house monitoring systems.
  • Analyse operational telemetry to identify systemic reliability concerns.
DevOps & CI/CD Enablement
  • Collaborate with engineering teams to improve CI/CD pipelines.
  • Implement deployment strategies (blue-green, canary, rolling updates).
  • Advocate for reliability-focused design patterns.
Security & Compliance
  • Ensure infrastructure adheres to security standards and compliance requirements.
  • Participate in vulnerability assessments and remediation.
Required Technical Skills
  • Strong production support and operations experience in enterprise environments.
  • Substantial experience leading Production Support Engineers, in a senior/lead or team-management capacity.
  • Strong experience with cloud platforms (AWS, Azure, and GCP).
  • Substantial, practical hands-on experience with both Windows and Linux operating systems, and sound working familiarity with Databricks and Microsoft Fabric for analytics workloads.
  • Expertise in monitoring & observability tools (e.g., Prometheus, Grafana, Datadog, or in-house tools).
  • Working knowledge of:
  • ServiceNow
  • Incident workflows,
  • Escalation management,
  • Operational support models.
  • Exposure to database technologies including:
  • Oracle Database
  • Microsoft SQL Server
  • MySQL
  • PostgreSQL
  • NoSQL ecosystems preferred.
  • Strong understanding of:
  • Windows and Linux systems,
  • Infrastructure monitoring,
  • Alerting concepts,
  • Production operations.
  • Experience supporting 24x7 enterprise production environments.
Preferred Qualifications
  • Experience working with in-house monitoring or observability product teams.
  • Familiarity with SRE/DevOps operational practices.
  • Exposure to enterprise event management systems.
  • Knowledge of automation/scripting (Python, Shell, PowerShell).
  • Experience handling high-severity production incidents.
Critical Non-Technical Skills
Operational Intuition
  • Ability to detect operational anomalies early.
  • Strong troubleshooting instinct and pattern recognition.
Fearless Communication
  • Ability to speak confidently during incidents and escalations.
  • Comfortable engaging customers and senior stakeholders of in-house management, along with multiple technical teams.
  • Exceptional written and verbal communication skills, able to present operational status and risk directly to customers and senior leadership with clarity and confidence.
Cross-Team Collaboration
  • Ability to coordinate effectively across DBA teams, support organizations, and development groups.
Calmness Under Pressure
  • Structured decision-making during high-severity incidents.
Ownership Mindset
  • Drives issues to closure rather than relying solely on assigned ownership boundaries.
Investigative Curiosity
  • Continuously analyses why operational failures occur and how they can be prevented.
  • Substantial track record leading, mentoring, and developing Production Support Engineers.
  • Sets the standard for operational excellence and coaches junior/mid-level engineers toward it.
  • Acts as an escalation point and mentor for less experienced engineers during high-severity incidents.
  • Team Leadership & Mentorship
Preferred Qualifications
  • Relevant Certifications in cloud platforms (AWS/Azure/GCP).
  • Familiarity with SRE/DevOps operational practices.
  • Familiarity with Databricks and Fabric domain.
  • Exposure to enterprise event management systems.
  • Knowledge of automation/scripting (Python, Shell, PowerShell).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Product Support Engineer
Product Support Engineer

Datavail Corp. • Mumbai

On-site
INR 2,800,000 - 6,000,000
Production Support Engineer
Production Support Engineer

Moofwd • Pune District

On-site
INR 600,000 - 1,000,000
Application Development Associate Director (Production Support)
Application Development Associate Director (Production Support)

Evernorth Health Services • Hyderabad

On-site
INR 3,500,000 - 6,000,000
Senior New Relic & AWS Cloud Operations Engineer
Senior New Relic & AWS Cloud Operations Engineer

Tata Consultancy Services • Kolkata District, Hyderabad, Chennai District

On-site
INR 3,500,000 - 6,000,000
Senior Product Support Engineer
Senior Product Support Engineer

SISA • Bengaluru

Hybrid
INR 1,200,000 - 2,100,000
Application Support Engineer
Application Support Engineer

Orcapod Consulting Services • Bengaluru

On-site
INR 1,500,000 - 2,200,000
Sr. SQL Server DBA Engineer
Sr. SQL Server DBA Engineer

Global Software Solutions Group • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Specialist Observability Engineering
Specialist Observability Engineering

Pearson • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Sr. SQL Server DBA Engineer
Sr. SQL Server DBA Engineer

GSSTech Group • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Applications Support Specialist
Applications Support Specialist

Ensono • Pune District

Hybrid
INR 2,600,000 - 4,200,000
Hybrid work model
On-call rotation