Senior Technical Support Specialist

HCL Technologies Limited

Dadri

Presencial

INR 1.800.000 - 3.000.000

Jornada completa

hace 11 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Convierte este puesto en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

HCLTech is seeking a GenAI & ML Operations Engineer to ensure reliability and governance of production GenAI/ML solutions. You will monitor AI workloads, resolve incidents, and implement observability while supporting model upgrades and lifecycle activities.

You will collaborate with Data Scientists, ML Engineers, Platform Engineers, and stakeholders to maintain scalable, well-governed AI services in production, with a focus on reliability, efficiency, and ongoing improvement.

Formación

  • Experience with GenAI/LLM platforms and AI observability.
  • Strong understanding of ML lifecycle and MLOps concepts.
  • Proficient in monitoring/observability tools like Datadog or Dynatrace.
  • Experience with cloud platforms (GCP) and containerized workloads (Kubernetes).
  • Scripting in Python and automation.

Responsabilidades

  • Ensure reliability, availability, and observability of production GenAI/ML solutions.
  • Design and maintain AI/ML observability dashboards, alerts, and scorecards.
  • Lead incident response, RCA, and preventive actions for production incidents.
  • Support model upgrades, retraining, and post-deployment validation.
  • Maintain runbooks, SOPs, and governance for security and Responsible AI.

Conocimientos

Generative AI/LLMs/RAG
MLOps
Observability/Monitoring
Cloud (GCP) & Kubernetes
Python scripting

Herramientas

Datadog
Dynatrace
Kubernetes
MemoryStore

Descripción del empleo

The GenAI & ML Operations Engineer is responsible for ensuring the reliability, availability, observability, and operational support of Production Generative AI and Machine Learning solutions. The role focuses on monitoring AI/ML platforms, resolving incidents, implementing observability, supporting model upgrades, executing minor enhancements, and driving operational excellence across AI products and services.

The engineer works closely with Data Scientists, ML Engineers, Platform Engineers, Product Teams, and Business Stakeholders to maintain stable, scalable, and well-governed AI solutions in production.

Key Responsibilities

AI/ML Production Operations

  • Monitor and support production GenAI and ML workloads, Perform daily operational health checks and service validations.
  • Troubleshoot and resolve operational issues impacting AI services, Support L1/L2/L3 incident management activities.
  • Ensure SLA, availability, and performance objectives are met.

Observability & Monitoring

  • Design, implement, and maintain AI/ML observability solutions, Create dashboards, reports, and operational scorecards.
  • Configure alerts for failures, threshold breaches, latency, data quality, model performance, and infrastructure issues.
  • Monitor model health, drift, performance, usage, and cost metrics, continuously improve monitoring coverage and alert effectiveness.

Incident & Problem Management

  • Respond to production incidents and service disruption, Conduct root cause analysis (RCA) and document findings.
  • Drive preventive actions to reduce recurring incidents, Participate in major incident support and service restoration activities.

Model Lifecycle Support

  • Support AI/ML model upgrades, version management, Monitor production behaviour following model changes and perform post-deployment validation.
  • Support retraining, tuning, and model refresh initiatives.
  • Deliver minor enhancements to AI applications and operational tooling, Automate repetitive operational tasks where feasible.
  • Identify opportunities to improve reliability, efficiency, and supportability, Contribute to AI Ops, MLOps, and operational maturity initiatives.

Documentation & Governance

  • Maintain runbooks, SOPs, knowledge articles, and support documentation, Document incidents, problems, technical debt, risks, and remediation plans.
  • Support change management and release governance processes, Ensure operational compliance with security and Responsible AI requirements.

Required Skills & Experience

Technical Skills

  • Generative AI, LLMs, RAG, and AI application support.
  • Machine Learning lifecycle and MLOps concepts.
  • Monitoring and Observability Platforms (Datadog, Dynatrace.).
  • Cloud Platforms (GCP), Kubernetes and containerized workloads, MemoryStore, Tracing, Storage, GAR.
  • Python and automation scripting.

Operational Skills

  • Incident Management, Problem Management, Change Management
  • Service Reliability Engineering (SRE)
  • Operational Reporting and Governance, Stakeholder Communication

Preferred Experience

  • 8+ years in IT Operations, DevOps, SRE, Cloud Operations, MLOps, or AI Operations.
  • 3+ years of experience with GenAI / LLM platforms / ML Models
  • Experience in working with Databricks jobs & pipelines, Catalog, compute, models, SQL warehouse.
  • Experience supporting production AI/ML solutions, working knowledge of ITIL-based operational processes
  • Familiarity with GenAI platforms, LLM-based applications, and AI observability practices.
Key Responsibilities

GenAI & ML Operations Engineer

Role Summary

The GenAI & ML Operations Engineer is responsible for ensuring the reliability, availability, observability, and operational support of Production Generative AI and Machine Learning solutions. The role focuses on monitoring AI/ML platforms, resolving incidents, implementing observability, supporting model upgrades, executing minor enhancements, and driving operational excellence across AI products and services.

The engineer works closely with Data Scientists, ML Engineers, Platform Engineers, Product Teams, and Business Stakeholders to maintain stable, scalable, and well-governed AI solutions in production.

Key Responsibilities

AI/ML Production Operations

  • Monitor and support production GenAI and ML workloads, Perform daily operational health checks and service validations.
  • Troubleshoot and resolve operational issues impacting AI services, Support L1/L2/L3 incident management activities.
  • Ensure SLA, availability, and performance objectives are met.

Observability & Monitoring

  • Design, implement, and maintain AI/ML observability solutions, Create dashboards, reports, and operational scorecards.
  • Configure alerts for failures, threshold breaches, latency, data quality, model performance, and infrastructure issues.
  • Monitor model health, drift, performance, usage, and cost metrics, continuously improve monitoring coverage and alert effectiveness.

Incident & Problem Management

  • Respond to production incidents and service disruption, Conduct root cause analysis (RCA) and document findings.
  • Drive preventive actions to reduce recurring incidents, Participate in major incident support and service restoration activities.

Model Lifecycle Support

  • Support AI/ML model upgrades, version management, Monitor production behaviour following model changes and perform post-deployment validation.
  • Support retraining, tuning, and model refresh initiatives.
  • Deliver minor enhancements to AI applications and operational tooling, Automate repetitive operational tasks where feasible.
  • Identify opportunities to improve reliability, efficiency, and supportability, Contribute to AI Ops, MLOps, and operational maturity initiatives.

Documentation & Governance

  • Maintain runbooks, SOPs, knowledge articles, and support documentation, Document incidents, problems, technical debt, risks, and remediation plans.
  • Support change management and release governance processes, Ensure operational compliance with security and Responsible AI requirements.

Required Skills & Experience

Technical Skills

  • Generative AI, LLMs, RAG, and AI application support.
  • Machine Learning lifecycle and MLOps concepts.
  • Monitoring and Observability Platforms (Datadog, Dynatrace.).
  • Cloud Platforms (GCP), Kubernetes and containerized workloads, MemoryStore, Tracing, Storage, GAR.
  • Python and automation scripting.

Operational Skills

  • Incident Management, Problem Management, Change Management
  • Service Reliability Engineering (SRE)
  • Operational Reporting and Governance, Stakeholder Communication

Preferred Experience

  • 8+ years in IT Operations, DevOps, SRE, Cloud Operations, MLOps, or AI Operations.
  • 3+ years of experience with GenAI / LLM platforms / ML Models
  • Experience in working with Databricks jobs & pipelines, Catalog, compute, models, SQL warehouse.
  • Experience supporting production AI/ML solutions, working knowledge of ITIL-based operational processes
  • Familiarity with GenAI platforms, LLM-based applications, and AI observability practices.
Skill Requirements

GenAI & ML Operations Engineer

Role Summary

The GenAI & ML Operations Engineer is responsible for ensuring the reliability, availability, observability, and operational support of Production Generative AI and Machine Learning solutions. The role focuses on monitoring AI/ML platforms, resolving incidents, implementing observability, supporting model upgrades, executing minor enhancements, and driving operational excellence across AI products and services.

The engineer works closely with Data Scientists, ML Engineers, Platform Engineers, Product Teams, and Business Stakeholders to maintain stable, scalable, and well-governed AI solutions in production.

Key Responsibilities

AI/ML Production Operations

  • Monitor and support production GenAI and ML workloads, Perform daily operational health checks and service validations.
  • Troubleshoot and resolve operational issues impacting AI services, Support L1/L2/L3 incident management activities.
  • Ensure SLA, availability, and performance objectives are met.

Observability & Monitoring

  • Design, implement, and maintain AI/ML observability solutions, Create dashboards, reports, and operational scorecards.
  • Configure alerts for failures, threshold breaches, latency, data quality, model performance, and infrastructure issues.
  • Monitor model health, drift, performance, usage, and cost metrics, continuously improve monitoring coverage and alert effectiveness.

Incident & Problem Management

  • Respond to production incidents and service disruption, Conduct root cause analysis (RCA) and document findings.
  • Drive preventive actions to reduce recurring incidents, Participate in major incident support and service restoration activities.

Model Lifecycle Support

  • Support AI/ML model upgrades, version management, Monitor production behaviour following model changes and perform post-deployment validation.
  • Support retraining, tuning, and model refresh initiatives.
  • Deliver minor enhancements to AI applications and operational tooling, Automate repetitive operational tasks where feasible.
  • Identify opportunities to improve reliability, efficiency, and supportability, Contribute to AI Ops, MLOps, and operational maturity initiatives.

Documentation & Governance

  • Maintain runbooks, SOPs, knowledge articles, and support documentation, Document incidents, problems, technical debt, risks, and remediation plans.
  • Support change management and release governance processes, Ensure operational compliance with security and Responsible AI requirements.

Required Skills & Experience

Technical Skills

  • Generative AI, LLMs, RAG, and AI application support.
  • Machine Learning lifecycle and MLOps concepts.
  • Monitoring and Observability Platforms (Datadog, Dynatrace.).
  • Cloud Platforms (GCP), Kubernetes and containerized workloads, MemoryStore, Tracing, Storage, GAR.
  • Python and automation scripting.

Operational Skills

  • Incident Management, Problem Management, Change Management
  • Service Reliability Engineering (SRE)
  • Operational Reporting and Governance, Stakeholder Communication

Preferred Experience

  • 8+ years in IT Operations, DevOps, SRE, Cloud Operations, MLOps, or AI Operations.
  • 3+ years of experience with GenAI / LLM platforms / ML Models
  • Experience in working with Databricks jobs & pipelines, Catalog, compute, models, SQL warehouse.
  • Experience supporting production AI/ML solutions, working knowledge of ITIL-based operational processes
  • Familiarity with GenAI platforms, LLM-based applications, and AI observability practices.

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026totaled $14.8billion.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Track Manager - Apache Kafka, Windows PowerShell
Track Manager - Apache Kafka, Windows PowerShell

HCL Technologies Limited • Bengaluru

Presencial
INR 5.000.000 - 7.500.000
Senior ML Ops Technical Specialist - DevOps, Python
Senior ML Ops Technical Specialist - DevOps, Python

HCL Technologies Limited • Chennai District

Presencial
INR 350.000 - 600.000
Senior Development Lead - Advanced RPA
Senior Development Lead - Advanced RPA

HCL Technologies Limited • Dadri

Presencial
INR 2.500.000 - 4.200.000
Senior RPA Developer - Core
Senior RPA Developer - Core

HCL Technologies Limited • India

Presencial
INR 1.200.000 - 2.200.000
GLO AI-ML specialist
GLO AI-ML specialist

Hewlett Packard Enterprise • Bengaluru

Híbrido
INR 1.800.000 - 3.200.000
Senior Project Lead
Senior Project Lead

HCL Technologies Limited • Hyderabad

Presencial
INR 1.800.000 - 2.600.000
Technical Specialist
Technical Specialist

HCL Technologies Limited • Dadri

Presencial
INR 1.500.000 - 2.100.000
DevOps and MLOps Engineer - Automation, Analytics and Generative AI
DevOps and MLOps Engineer - Automation, Analytics and Generative AI

Hewlett Packard Enterprise India Private Limited • Bengaluru

Presencial
INR 4.000.000 - 7.000.000
Health & Wellbeing
Personal & Professional Development
Unconditional Inclusion
Campus-Competent
Campus-Competent

HCL Technologies Limited • Gautam Buddha Nagar

Presencial
INR 1.200.000 - 1.800.000
Senior Group Technical Architect
Senior Group Technical Architect

HCL Technologies Limited • India

Presencial
INR 1.800.000 - 3.000.000