Job Title: Engineer - AWS, Azure, GCP, AKS
Experience: 6+ Years
Education: Any Graduate
Location: Mumbai
Job Description
Role Overview
- The in-house monitoring product development and QE team(s),
- Internal ITSM/ticketing teams using ServiceNow,
- Multiple enterprise DBA organizations across:
- Oracle Database
- Open-source database platforms (MySQL, PostgreSQL, MongoDB, Cassandra, etc.)
- Microsoft SQL Server
Key Responsibilities
Production Operations & Monitoring Support
- Provide operational ownership and production support for the in-house enterprise monitoring platform.
- Monitor health, performance, alerting quality, and operational stability of monitoring services.
- Analyze monitoring gaps, false positives, missed alerts, and operational inefficiencies.
- Ensure monitoring coverage across Oracle, Open-source, and SQL Server database environments.
Incident & Escalation Management
- Act as the operational point-of-contact during production incidents involving monitoring failures, alerting gaps, or infrastructure issues.
- Coordinate incident triage across DBA teams, monitoring development teams, infrastructure teams, service management teams.
- Drive bridge calls and ensure effective stakeholder communication during critical outages.
- Perform root cause analysis (RCA) and post-incident operational reviews.
ServiceNow & Ticket Workflow Coordination
- Work with ServiceNow for operational escalations, service requests.
- Review ticket quality and ensure operational accuracy of issue classification and routing.
- Improve ticket workflows between DBA teams and monitoring platform support teams.
- Collaborate with internal support organizations to streamline escalation processes.
Cross-Functional DBA Collaboration
- Collaborate closely with enterprise DBA teams supporting Oracle Database, MySQL, PostgreSQL, MongoDB, Apache Cassandra, Microsoft SQL Server, Cloud services (AWS, AZURE, GCP).
- Understand operational monitoring requirements specific to each database technology.
- Work with DBAs to validate alert thresholds, event correlation, and monitoring accuracy.
- Serve as the operational liaison between DBAs and monitoring team developers and QE.
Operational Excellence & Reliability Engineering
- Identify recurring operational pain points and recommend automation opportunities.
- Improve alert quality, event correlation, and monitoring reliability.
- Participate in operational readiness reviews for new monitoring features.
- Help define operational standards, playbooks, and escalation procedures.
Monitoring & Observability Engineering
- Support enterprise observability initiatives involving metrics, events, alerting, dashboards, health monitoring, incident correlation.
- Work with both commercial and in-house monitoring systems.
- Analyse operational telemetry to identify systemic reliability concerns.
DevOps & CI/CD Enablement
- Collaborate with engineering teams to improve CI/CD pipelines.
- Implement deployment strategies (blue-green, canary, rolling updates).
- Advocate for reliability-focused design patterns.
Security & Compliance
- Ensure infrastructure adheres to security standards and compliance requirements.
- Participate in vulnerability assessments and remediation.
Required Technical Skills
- Strong production support and operations experience in enterprise environments.
- Strong experience with cloud platforms (AWS, Azure, or GCP).
- Expertise in monitoring & observability tools (e.g., Prometheus, Grafana, Datadog, or in-house tools).
- Working knowledge of ServiceNow, incident workflows, escalation management, operational support models.
- Exposure to database technologies including Oracle Database, Microsoft SQL Server, MySQL, PostgreSQL, NoSQL ecosystems.
- Strong understanding of Linux systems, infrastructure monitoring, alerting concepts, production operations.
- Experience supporting 24x7 enterprise production environments.
Preferred Qualifications
- Experience working with in-house monitoring or observability product teams.
- Familiarity with SRE/DevOps operational practices.
- Exposure to enterprise event management systems.
- Knowledge of automation/scripting (Python, Shell, PowerShell).
- Experience handling high-severity production incidents.
- Certifications in cloud platforms (AWS/Azure/GCP).
Critical Non-Technical Skills
An ideal candidate must demonstrate:
Operational Intuition
- Ability to detect operational anomalies early.
- Strong troubleshooting instinct and pattern recognition.
Fearless Communication
- Ability to speak confidently during incidents and escalations.
- Comfortable engaging senior stakeholders and multiple technical teams.
Cross-Team Collaboration
- Ability to coordinate effectively across DBA teams, support organizations, and development groups.
Calmness Under Pressure
- Structured decision-making during high-severity incidents.
Ownership Mindset
- Drives issues to closure rather than relying solely on assigned ownership boundaries.
Investigative Curiosity
- Continuously analyses why operational failures occur and how they can be prevented.