Senior Technical Lead

HCL Technologies Limited

Pune District

On-site

INR 4,000,000 - 6,000,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

HCLTech is seeking a Lead/Manager of AI Observability to own how we see, measure, and trust our AI systems in production. This role sits at the intersection of MLOps, platform engineering, and applied AI, building monitoring, tracing, evaluation, and alerting infrastructure for real-time visibility into model behavior, cost, and risk.

You will define observability strategy for LLM-powered and traditional ML systems, lead a small team of engineers, and partner with ML/platform reliability teams

Qualifications

  • + years in software/ML engineering with 3+ years on observability, monitoring, reliability, or MLOps.
  • Hands-on experience running AI/ML systems in production, including at least one large model.
  • Strong understanding of tracing, logging, and metrics pipelines in AI workflows.

Responsibilities

  • Own the end-to-end observability strategy for AI systems in production, including tracing, logging, metrics, evaluation pipelines, and alerting.
  • Design and build monitoring for model/agent quality: drift, hallucination, latency, cost per request, token usage, task success rate.
  • Establish golden signals and SLOs distinct from infra SLOs (output quality, safety, groundedness, factuality).
  • Build/integrate tracing across multi-step/agentic workflows to root-cause across prompts and model versions.
  • Stand up automated evaluation frameworks to score live traffic and catch regressions after updates.
  • Partner with ML/platform eng to instrument models with observability hooks before production.
  • Define incident response for AI failures (silent quality degradation, drift, prompt injection, unsafe outputs).
  • Build dashboards and reports for engineering, product, and leadership on AI health, cost, and risk.
  • Lead or mentor a team of observability engineers; act as technical lead across ML/platform teams.
  • Evaluate and manage observability toolchain (build vs buy) across tracing, evals, monitoring, cost-tracking.

Skills

Observability
MLOps
AI systems in prod
Distributed tracing
Software engineering
Python/Go

Education

Bachelor's in CS or related

Tools

OpenTelemetry
Vector databases

Job description

We're looking for a Lead / Manager of AI Observability to own how we see, measure, and trust our AI systems once they're live in production. This role sits at the intersection of MLOps, platform engineering, and applied AI, and is responsible for building the monitoring, tracing, evaluation, and alerting infrastructure that gives engineering, product, and leadership real-time visibility into model and agent behavior, quality, cost, and risk. You'll define the observability strategy for LLM-powered and traditional ML systems in production, lead a small team (or embedded function) of engineers, and partner closely with ML, platform, and reliability teams to catch issues before customers do.

Key Responsibilities
  • Own the end-to-end observability strategy for AI systems in production, including tracing, logging, metrics, evaluation pipelines, and alerting.
  • Design and build systems to monitor model/agent quality in production: accuracy drift, hallucination rate, latency, cost per request, token usage, and task success rate.
  • Establish golden signals and SLOs for AI systems, distinct from traditional infra SLOs (e.g., output quality, safety, groundedness, factuality).
  • Build or integrate tracing across multi-step/agentic workflows so failures can be root-caused across prompts, tool calls, retrieval steps, and model versions.
  • Stand up automated evaluation frameworks (offline and online/production evals) to continuously score live traffic and catch regressions after model, prompt, or data updates.
  • Partner with ML/platform engineering to instrument new models and features with observability hooks before they reach production.
  • Define and drive incident response processes specific to AI failures (silent quality degradation, drift, prompt injection, unsafe outputs) — not just uptime.
  • Build dashboards and reporting for engineering, product, and executive stakeholders on production AI health, cost, and risk posture.
  • Lead, mentor, and grow a team of engineers focused on observability tooling, or act as the technical lead embedded across ML/platform teams.
  • Evaluate, select, and manage the observability toolchain (build vs. buy) across tracing, evals, monitoring, and cost-tracking platforms.
  • Partner with security, compliance, and legal on auditability, data retention, and responsible-AI monitoring requirements.
Skill Requirements
  • + years in software/ML engineering, with 3+ years focused on observability, monitoring, reliability, or MLOps.
  • Hands-on experience running AI/ML systems in production, including at least one LLM-based or generative AI system at scale.
  • Strong understanding of distributed tracing, structured logging, and metrics pipelines (e.g., OpenTelemetry-style concepts), applied to AI/agentic workflows.
  • Experience designing evaluation frameworks for generative AI (offline benchmarks, online/production evals, human-in-the-loop review).
  • Solid grasp of the unique failure modes of production AI: drift, hallucination, prompt injection, latency/cost blowups, silent quality regressions.
  • Track record of building or leading a team, or serving as a technical lead across cross-functional engineering groups.
  • Proficiency in at least one major programming language (Python, Go, or similar) and comfort working across the ML/platform stack.
  • Excellent cross-functional communication — able to translate observability data into decisions for engineers, product managers, and executives.
Other Requirements
  • Experience with vector databases, RAG pipelines, or agentic frameworks in production.
  • Background in SRE/DevOps prior to moving into ML/AI observability.
  • Familiarity with responsible AI, model risk management, or AI governance frameworks.
  • Experience presenting observability/risk posture to executive or board-level audiences.

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026totaled $14.8billion.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Solution Architect I
AI Solution Architect I

HCL Technologies Limited • Dadri

On-site
INR 3,200,000 - 7,600,000
Sr Tower Lead (Tools & Automation)
Sr Tower Lead (Tools & Automation)

HCL Technologies Limited • Dadri

On-site
INR 1,800,000 - 3,200,000
Platform Engineer II
Platform Engineer II

HCL Technologies Limited • India

On-site
INR 1,800,000 - 3,200,000
Senior Technical Specialist
Senior Technical Specialist

HCL Technologies Limited • Bengaluru

Hybrid
INR 2,500,000 - 3,500,000
AI Operations Engineering Technical Leader
AI Operations Engineering Technical Leader

Cisco • Bengaluru

On-site
INR 3,500,000 - 5,500,000
AI Solution Principal I
AI Solution Principal I

HCL Technologies Limited • Dadri, Bengaluru

On-site
INR 3,500,000 - 6,000,000
AI Solutions Engineer I (Tech)
AI Solutions Engineer I (Tech)

Hyscaler • Khordha

On-site
INR 800,000 - 1,200,000
Continuous learning opportunities
Inclusive company culture
Team-centric activities
+2
Staff Ml Ops Engineer [T500-28545]
Staff Ml Ops Engineer [T500-28545]

ANSR • Bengaluru

On-site
INR 4,500,000 - 7,000,000
AI Engineering Lead
AI Engineering Lead

Blend360 • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Senior AI and Machine Learning Engineer
Senior AI and Machine Learning Engineer

Hewlett Packard Enterprise Development LP • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Health & Wellbeing benefits
Personal & Professional Development programs
Inclusive work culture