Observability Technical Lead

Long Business Systems, Inc.

Chennai District

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Long Business Systems, Inc. is seeking a senior Observability Lead to direct a team responsible for monitoring, logging, tracing, and operational visibility across critical applications and infrastructure.

You will drive reliability, incident readiness, and continuous improvement while partnering with SREs, GCC and Security teams. You will define the observability strategy, oversee dashboards and alerting, and lead AI-powered enhancements including Copilot-enabled workflows to automate

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field.
  • 8+ years in infrastructure, operations, SRE, platform engineering, or observability.
  • 3+ years of people management leading technical teams.
  • Strong knowledge of observability platforms (Splunk, Datadog, AppDynamics, Dynatrace, Grafana, Prometheus, OpenTelemetry).
  • Experience in cloud environments (Azure & GCP).

Responsibilities

  • Lead and mentor observability engineers supporting enterprise platforms and services.
  • Define and execute the observability strategy, standards, and roadmap.
  • Oversee monitoring, logging, alerting, tracing, and dashboarding solutions.
  • Drive service reliability, incident response readiness, and operational excellence initiatives.
  • Collaborate with application, infrastructure, cloud, and SRE teams to improve system health and performance.
  • Establish KPIs, SLAs, and metrics to measure platform reliability and team effectiveness.
  • Manage hiring, performance development, resource planning, and stakeholder communications.
  • Ensure adoption of best practices for observability, automation, and proactive problem management.
  • Champion AI and Copilot-enabled workflows within Observability.
  • Evaluate and implement AI-driven monitoring, alert correlation, and incident management capabilities.
  • Define and execute AI strategy for intelligent automation in observability and operations ecosystem.
  • Drive adoption of AI-powered operations including autonomous incident management and predictive analytics.
  • Lead initiatives leveraging Copilot, Agentic AI frameworks, and AI agents to automate workflows and resilience improvements.
  • Partner with engineering to identify use cases for Agentic AI to reduce manual effort and improve efficiency.
  • Partner with engineering and platform teams to build intelligent dashboards and automated remediation solutions.

Skills

Leadership
Stakeholder management
Communication
Problem solving

Education

Bachelor's degree in CS/Eng

Tools

Splunk
Datadog
AppDynamics
Dynatrace
Grafana
Prometheus
OpenTelemetry
Azure
GCP

Job description

This role will lead a team responsible for monitoring, logging, tracing, and operational visibility across critical applications and infrastructure. The candidate will drive operational excellence, reliability, and continuous improvement while partnering closely with SREs, GCC and Security teams.

Key Responsibilities
  • Lead and mentor a team of observability engineers supporting enterprise platforms and services.
  • Define and execute the observability strategy, standards, and roadmap.
  • Oversee monitoring, logging, alerting, tracing, and dashboarding solutions.
  • Drive service reliability, incident response readiness, and operational excellence initiatives.
  • Collaborate with application, infrastructure, cloud, and SRE teams to improve system health and performance.
  • Establish KPIs, SLAs, and operational metrics to measure platform reliability and team effectiveness.
  • Manage hiring, performance development, resource planning, and stakeholder communications.
  • Ensure adoption of best practices for observability, automation, and proactive problem management.
  • Champion the adoption of AI and Copilot-enabled workflows within the Observability organization.
  • Evaluate and implement AI-driven monitoring, alert correlation, and incident management capabilities.
  • Define and execute the organization's strategy for AI, Agentic AI, and intelligent automation within the observability and operations ecosystem.
  • Drive the adoption of AI-powered operations, including autonomous incident management, intelligent alert correlation, predictive analytics, and self-healing platforms.
  • Lead initiatives leveraging Microsoft Copilot, Agentic AI frameworks, and AI agents to automate operational workflows, knowledge management, problem resolution, and service reliability improvements.
  • Partner with engineering teams to identify and prioritize use cases for Agentic AI that reduce manual effort and improve operational efficiency.
  • Partner with engineering and platform teams to build intelligent operational dashboards and automated remediation solutions.
Qualifications
  • Bachelor's degree in Computer Science, Engineering, or a related field.
  • 8+ years of experience in infrastructure, operations, SRE, platform engineering, or observability domains.
  • 3+ years of people management experience leading technical teams.
  • Strong knowledge of observability platforms such as Splunk, Datadog, AppDynamics, Dynatrace, Grafana, Prometheus, OpenTelemetry, or similar tools.
  • Experience working in cloud environments (Azure & GCP).

Experience leveraging Microsoft Copilot, Generative AI, and AI-powered observability capabilities to improve operational efficiency, incident response, and engineering productivity.

Knowledge of AI-assisted troubleshooting, anomaly detection, root cause analysis, and predictive monitoring solutions.

  • Excellent communication, stakeholder management, and leadership skills.
Preferred
  • Experience leading globally distributed teams.
  • Strong background in automation, DevOps, and reliability engineering practices.
  • Familiarity with enterprise-scale monitoring and incident management processes.
  • Experience with Generative AI, Agentic AI, Microsoft Copilot, Azure AI, OpenAI technologies, or similar AI platforms.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
AI Observability Principal Architect
AI Observability Principal Architect

LTM • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Hyderabad

On-site
INR 4,200,000 - 7,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Chennai District

On-site
INR 3,500,000 - 7,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Mumbai

On-site
INR 4,000,000 - 6,000,000
AI Observability & Monitoring Engineer
AI Observability & Monitoring Engineer

Elabs Infotech • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Observability COE Lead- IT Consulting - Delhi NCR
Observability COE Lead- IT Consulting - Delhi NCR

Michael Page • Dadri

On-site
INR 4,000,000 - 7,000,000
Sony - Lead Platform Engineer - Observability Services
Sony - Lead Platform Engineer - Observability Services

Sony India Software Centre • Bengaluru

On-site
INR 1,200,000 - 1,500,000
Staff Observability Engineer
Staff Observability Engineer

I00M05 Wipro GE Healthcare Private Limited • Bengaluru

On-site
INR 1,500,000 - 2,500,000