Observability COE Lead- IT Consulting - Delhi NCR

Michael Page

Dadri

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Michael Page is seeking a Senior Leader to establish and grow an enterprise-wide Observability Centre of Excellence in a leading IT services environment. You will drive AI-powered observability across enterprise environments and shape the monitoring strategy for metrics, logs, traces, APM and user experience.

You will mentor a team of observability and reliability engineers, collaborate with engineering and business units, and influence cloud-native monitoring across AWS, Azure, Kubernetes, and

Qualifications

  • 12+ years in Observability, SRE, or infrastructure monitoring.
  • Proven experience building or scaling Observability CoE or reliability engineering capability.
  • Hands-on with Datadog, Dynatrace, New Relic, Grafana, Prometheus and OpenTelemetry.
  • Strong cloud knowledge (AWS/Azure) and Kubernetes.

Responsibilities

  • Lead establishment and growth of an enterprise-wide Observability CoE.
  • Drive adoption of observability best practices across infrastructure, cloud and application environments.
  • Own the monitoring strategy covering metrics, logs, traces, APM and user experience monitoring.
  • Develop and standardise observability frameworks, onboarding processes, governance and operating models.
  • Partner with engineering, platform and business teams to improve service reliability and operational resilience.
  • Define and implement SRE practices including SLI/SLO frameworks, error budgets and reliability metrics.
  • Drive automation, AIOps, anomaly detection and proactive incident management initiatives.
  • Build executive dashboards and translate operational insights into measurable business outcomes.
  • Mentor a specialist team of observability and reliability engineers while remaining hands-on with architecture and solution design.
  • Influence cloud-native monitoring strategies across AWS, Azure, Kubernetes and hybrid environments.

Skills

Observability leadership
SRE / Platform Engineering
Cloud expertise
Kubernetes
Docker
CI/CD pipelines
AIOps
Incident management
Stakeholder management

Tools

Datadog
Dynatrace
New Relic
Grafana
Prometheus
OpenTelemetry

Job description

  • Be recognised as a Senior Leader in a leading IT Services Organization
  • Chance to drive AI-powered observability across enterprise environments.
About Our Client

Our client is a large, globally established organisation with a strong focus on technology transformation, cloud adoption, infrastructure resilience, and operational excellence. The organisation is investing heavily in AI-driven observability, automation, and Site Reliability Engineering practices to improve service performance, customer experience, and business value.

Job Description
  • Lead the establishment and growth of an enterprise-wide Observability Centre of Excellence (CoE).
  • Drive adoption of observability best practices across infrastructure, cloud and application environments.
  • Own the monitoring strategy covering metrics, logs, traces, APM and user experience monitoring.
  • Develop and standardise observability frameworks, onboarding processes, governance and operating models.
  • Partner with engineering, platform and business teams to improve service reliability and operational resilience.
  • Define and implement SRE practices including SLI/SLO frameworks, error budgets and reliability metrics.
  • Drive automation, AIOps, anomaly detection and proactive incident management initiatives.
  • Build executive dashboards and translate operational insights into measurable business outcomes.
  • Mentor a specialist team of observability and reliability engineers while remaining hands-on with architecture and solution design.
  • Influence cloud-native monitoring strategies across AWS, Azure, Kubernetes and hybrid environments.
The Successful Applicant
  • 12+ years of experience in Observability, Site Reliability Engineering (SRE), Platform Engineering or Infrastructure Monitoring.
  • Proven experience building or scaling an Observability Practice, CoE or Reliability Engineering capability.
  • Strong hands-on expertise with Datadog, Dynatrace, New Relic, Grafana, Prometheus and OpenTelemetry.
  • Experience in application performance monitoring, distributed tracing, log analytics and telemetry instrumentation.
  • Strong knowledge of cloud platforms including AWS and Azure.
  • Experience with Kubernetes, Docker, CI/CD pipelines and modern engineering practices.
  • Exposure to AIOps, observability automation, anomaly detection and AI-enabled operations.
  • Deep understanding of incident management, RCA processes, reliability engineering and operational excellence.
  • Previous experience leading small, high-performing teams of engineers, architects or SMEs. (5-10 Member Teams)
  • Excellent stakeholder management and communication skills with the ability to engage senior leadership.
What\'s on Offer

Join a high-visibility role where you\'ll have the opportunity to build and shape an enterprise-wide Observability Centre of Excellence from the ground up. You\'ll influence technology strategy across cloud, infrastructure and application landscapes while driving modern SRE, AIOps and reliability engineering practices. This position offers strong exposure to executive stakeholders, ownership of observability transformation initiatives, and the chance to leverage cutting-edge AI-driven monitoring technologies.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Technical Lead
Observability Technical Lead

Jeevan Technologies • Chennai District

On-site
INR 3,000,000 - 6,000,000
Observability Technical Lead
Observability Technical Lead

Long Business Systems, Inc. • Chennai District

On-site
INR 4,000,000 - 7,000,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Site Reliability Engineer (SRE) / Observability Engineer
Site Reliability Engineer (SRE) / Observability Engineer

N Human Resources & Management Systems • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work
Certification reimbursement
Structured learning
Software Engineer-observability engineer
Software Engineer-observability engineer

Tranzeal • Bengaluru

Hybrid
INR 3,200,000 - 6,400,000
AI Observability & Monitoring Engineer
AI Observability & Monitoring Engineer

Elabs Infotech • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Enterprise Observability Platform Engineer
Enterprise Observability Platform Engineer

Be a Catalyst • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Kiya.ai - Observability Integration Lead
Kiya.ai - Observability Integration Lead

Infrasoft Technologies Ltd • Mumbai

On-site
INR 1,200,000 - 1,600,000
Observability Engineer (SRE)
Observability Engineer (SRE)

Cloudstepin • Hyderabad

On-site
INR 1,200,000 - 1,800,000
SRE Practice Lead
SRE Practice Lead

Coforge • Dadri

On-site
INR 3,000,000 - 4,200,000