AI Observability & Monitoring Engineer

Elabs Infotech

Bengaluru

On-site

INR 3,000,000 - 5,400,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Elabs Infotech in Bengaluru seeks an AI Observability Principal Architect to lead enterprise observability, SRE, and AIOps strategies. You will define telemetry standards, implement modern frameworks like Grafana and OpenTelemetry, and drive AI-driven monitoring across enterprise platforms.

The role requires architect-level expertise, strong leadership, and collaboration with ITSM and cross-functional teams to enable automated remediation and self-healing of critical systems.

Qualifications

  • 13+ years of experience in Observability, SRE, IT Operations, Platform Engineering, Monitoring, or AIOps.
  • Proven experience designing and implementing enterprise observability architectures.
  • Strong hands-on and architectural expertise in Grafana and OpenTelemetry.
  • Strong understanding of monitoring, telemetry, alerting, incident management, and automation.
  • Experience implementing SRE principles, including SLIs, SLOs, error budgets, and reliability engineering.
  • Experience integrating observability platforms with ITSM tools, preferably ServiceNow.
  • Strong leadership and stakeholder management across infra, app, SRE, and service management teams.

Responsibilities

  • Define and implement the enterprise observability strategy, architecture, standards, and best practices.
  • Design and implement observability solutions covering logs, metrics, and distributed traces using OpenTelemetry.
  • Develop enterprise dashboards, alerting frameworks, and operational views using Grafana.
  • Establish SLIs, SLOs, Error Budgets, and CUJs in collaboration with SRE and apps teams.
  • Lead AIOps and Event Management initiatives including correlation, anomaly detection, and predictive analytics.
  • Design and implement incident automation, intelligent routing, remediation, and self-healing capabilities.
  • Integrate observability platforms with ITSM platforms, preferably ServiceNow ITOM.
  • Drive automation using PowerShell and Ansible across environments.
  • Collaborate with SRE, Infra, App, Cloud, and Service Management teams.
  • Build executive-level dashboards for service health and reliability.
  • Provide technical leadership during major incidents and drive continuous service improvement.
  • Evaluate and implement AI/ML-driven observability and monitoring solutions.

Skills

SRE Practices
AIOps
Incident Management
AI/ML-driven Monitoring
Leadership

Tools

Grafana
OpenTelemetry
Splunk
SolarWinds
PowerShell
Ansible
ServiceNow ITOM

Job description

Experience: 13 - 16 Years

Location: Bellandur, Bangalore

Shift: India / US Shift

Role: Principal Architect

About the Role

We are looking for an experienced AI Observability Principal Architect to define and drive enterprise-wide observability, SRE, AIOps, and intelligent automation strategies. The ideal candidate will have strong expertise in Grafana, OpenTelemetry, SRE practices, AIOps, event management, automation, and infrastructure observability.

The role requires an architect who can establish observability standards, implement modern telemetry frameworks, drive AI-driven monitoring initiatives, and enable automated incident management and self-healing across enterprise platforms.

Key Responsibilities
  • Define and implement the enterprise observability strategy, architecture, standards, and best practices.
  • Design and implement observability solutions covering logs, metrics, and distributed traces using OpenTelemetry.
  • Develop enterprise dashboards, alerting frameworks, and operational views using Grafana.
  • Establish SLIs, SLOs, Error Budgets, and Critical User Journeys (CUJs) in collaboration with SRE and application teams.
  • Lead AIOps and Event Management initiatives including event correlation, anomaly detection, predictive analytics, and alert-noise reduction.
  • Design and implement incident automation, intelligent routing, remediation, and self-healing capabilities.
  • Integrate observability platforms with ITSM platforms, preferably ServiceNow ITOM.
  • Drive automation using PowerShell and Ansible.
  • Collaborate with SRE, Infrastructure, Application, Cloud, and Service Management teams.
  • Build executive-level and operational dashboards for service health, performance, availability, and reliability.
  • Provide technical leadership during major incidents and drive continuous service improvement.
  • Evaluate and implement AI/ML-driven observability and monitoring solutions.
  • Support automation initiatives using Power Apps and Power Automate where applicable.
Primary Skills
  • Grafana Observability, Dashboarding & Alerting
  • OpenTelemetry Logs, Metrics & Distributed Tracing
  • SRE Practices – SLI, SLO, Error Budgets, CUJs
  • AIOps & Event Management
  • Splunk / SolarWinds
  • Azure Observability / Azure Monitoring
  • Incident Management & Automation
  • PowerShell
  • Ansible
  • AI/ML-driven Monitoring and Predictive Analytics
  • Strong understanding of Windows/Linux infrastructure, networks, and firewalls
Secondary / Nice-to-Have Skills
  • ServiceNow ITOM
    • Service Mapping
    • IntegrationHub / integrations
    • Identification and Reconciliation Engine (IRE)
    • AIOps
  • ServiceNow FSM
  • ServiceNow ITAM / HAM
  • Power Apps
  • Power Automate
  • Cloud and platform observability
  • Predictive analytics
  • AI-driven monitoring and anomaly detection
Required Experience
  • 13+ years of experience across Observability, SRE, IT Operations, Platform Engineering, Monitoring, or AIOps.
  • Proven experience designing and implementing enterprise observability architectures.
  • Strong hands‑on and architectural expertise in Grafana and OpenTelemetry.
  • Strong understanding of monitoring, telemetry, alerting, incident management, and automation.
  • Experience implementing SRE principles, including SLIs, SLOs, error budgets, and reliability engineering.
  • Experience integrating observability platforms with ITSM tools, preferably ServiceNow.
  • Strong leadership and stakeholder management skills with the ability to work across infrastructure, application, SRE, and service management teams.
Preferred Profile

Candidates with experience in the following areas will be preferred:

  • Enterprise-scale AIOps transformation
  • AI/ML-based anomaly detection and predictive monitoring
  • ServiceNow ITOM / AIOps
  • Automated remediation and self-healing infrastructure
  • Azure observability
  • Large-scale Grafana and OpenTelemetry implementations
  • Major Incident Management and Continuous Service Improvement
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Observability Principal Architect
AI Observability Principal Architect

LTM • Bengaluru

On-site
INR 4,500,000 - 7,500,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Observability Technical Lead
Observability Technical Lead

Jeevan Technologies • Chennai District

On-site
INR 3,000,000 - 6,000,000
Site Reliability Engineer
Site Reliability Engineer

TerraGiG • India

On-site
INR 1,500,000 - 2,500,000
Kiya.ai - Observability Integration Lead
Kiya.ai - Observability Integration Lead

Infrasoft Technologies Ltd • Mumbai

On-site
INR 1,200,000 - 1,600,000
Observability Technical Lead
Observability Technical Lead

Long Business Systems, Inc. • Chennai District

On-site
INR 4,000,000 - 7,000,000
Senior Observability Engineer - Grafana & Prometheus
Senior Observability Engineer - Grafana & Prometheus

Zensar • Pune District, Bengaluru

Hybrid
INR 1,800,000 - 3,200,000
Software Engineer-observability engineer
Software Engineer-observability engineer

Tranzeal • Bengaluru

Hybrid
INR 3,200,000 - 6,400,000
Senior Engineer / Lead / Architect
Senior Engineer / Lead / Architect

Tranzeal • Hyderabad, Chennai District, Bengaluru

On-site
INR 4,500,000 - 7,000,000
Observability Engineer
Observability Engineer

Coforge • Hyderabad, Pune District, Greater Noida

Hybrid
INR 2,400,000 - 4,200,000