AIOps Leader

Premium Aerotec

Bengaluru

On-site

INR 3,000,000 - 7,000,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Premium Aerotec in Bengaluru seeks an AIOps Leader to architect and drive automation across our AI product support platform. You will bridge software engineering, SRE, and AI/ML domains to build self-healing infrastructure, predictive monitoring, and intelligent automation that reduces toil and MTTR.

You will lead cross-functional teams, design scalable pipelines, and implement RAG tooling, telemetry, and agent-based resolutions for Tier-1/2 support, aligning with enterprise observability and BI

Qualifications

  • Bachelor's or master's degree in CS/SE/IT or related field.
  • 7+ years of hands-on experience in software engineering, SRE, DevOps, or systems operations with focus on cloud automation and AI/ML tooling.
  • At least 3+ years architecting and running operations in major cloud ecosystems (AWS or GCP), including native AI/ML compute platforms.

Responsibilities

  • Self-Healing Infrastructure Workflows: design and deploy auto-remediation routines and event-driven automation.
  • LLM RAG Systems Engineering: build retrieval-augmented knowledge tools indexing telemetry and RCAs.
  • Bot Agent Development: develop conversational AI agents integrated into ticketing for Tier-1/2 resolutions.
  • Observability Architecture: create telemetry ingestion and automated log enrichment tied to incident tickets.
  • Predictive Anomaly Detection: real-time alerts and AI-driven monitoring across cloud services.
  • Operational BI Analytics: dashboards tracking MTTR, uptime, defect density, and SLA/CSAT compliance.
  • Toil Elimination: automate bottlenecks into autonomous production workflows.
  • Log Metadata Standardization: unify log outputs and tagging for AI engines.
  • Technical Mentorship: guide L1/L2 on automation and log parsing.

Skills

Software Engineering
Site Reliability Engineering (SRE)
DevOps
AI/ML tooling
Cloud automation

Education

Bachelor's or Master's degree in CS/SE/IT

Tools

Kubernetes (EKS/GKE)
Terraform
Jira/ServiceNow APIs
Datadog
OpenSearch
Vertex AI

Job description

Job Summary

Role: AIOps Leader

Job Description Date: August 2026

We are seeking a seasoned, hands-on AIOps Leader (7+ Years of Experience) to serve as the principal technical architect and lead builder for our centralized AI Product Support Operations function. Operating within a high-growth AI PSL (Product/Service Line), you will design, architect, and execute the end-to-end automation strategy that transforms raw operational chaos into scalable, self-healing, and data-driven infrastructure.

In this role, you will bridge software engineering, site reliability, and AI/ML architectures. You will lead the creation of intelligent diagnostic pipelines, custom RAG-driven knowledge tools, self-healing systems, and automated triage engines. You will work closely with cross-functional leadership, L1/L2 support teams, and platform engineering to systematically eliminate operational toil, optimize MTTR, and build proactive anomaly detection mechanisms across our AI ecosystem.

Number of positions: 1

Qualifications
  • Education: Bachelor s or Master s degree in Computer Science, Software Engineering, Information Technology, or a related quantitative field.
  • Overall Experience: 7+ years of hands-on experience across Software Engineering, Site Reliability Engineering (SRE), DevOps, or Systems Operations-with a focused concentration on cloud infrastructure automation and AI/ML operational tooling.
  • Platform Specialization: At least 3+ years architecting and running operations directly within major cloud ecosystems (AWS or GCP), including native AI/ML compute platforms.
Responsibilities
  • Architecture Advanced AI/ML Automation
    • Self-Healing Infrastructure Workflows: Architect, build, and deploy auto-remediation routines, script-based diagnostic runners, and event-driven automation triggers that autonomously resolve platform issues.
    • LLM RAG Systems Engineering: Design, implement, and maintain advanced Retrieval-Augmented Generation (RAG) knowledge tools, vector databases, and LLM utilities that index telemetry, historic logs, and RCAs for instant incident context.
    • Bot Agent Development: Lead the development and production rollout of conversational AI agents, custom webhooks, and self-service bots integrated into ticketing engines to automate Tier-1 and Tier-2 resolutions.
  • Observability, Telemetry Predictive Analytics
    • Observability Architecture: Build enterprise-grade telemetry ingestion workflows, automated log scraping, and context-enrichment pipelines that dynamically append system metrics directly to incident tickets upon creation.
    • Predictive Anomaly Detection: Configure and tune real-time predictive alerting, log-pattern analysis, and AI-driven monitoring models across AWS, Azure, or GCP microservices.
    • Operational BI Analytics: Architect and own centralized executive and operational dashboards (e.g., ServiceNow, Datadog) tracking MTTR velocity, system uptime, defect density, ticket deflection rates, and SLA/CSAT compliance.
  • Operational Engineering L1/L2 Empowerment
    • Toil Elimination: Continuously audit support operational bottlenecks across product teams, transforming high-volume manual intervention points into production-grade, single-click, or fully autonomous workflows.
    • Log Metadata Standardization: Standardize system log outputs, stack-trace formatting, and tagging taxonomy across all AI products to ensure platform telemetry remains machine-readable for AI engines.
    • Technical Mentorship: Guide L1/L2 support engineers on best practices for automation, code-based triage, and log parsing.
Technical Essentials
  • AWS GCP Native AI Platforms: Deep hands-on experience orchestrating production AI/ML workflows on AWS (Bedrock, SageMaker AI, OpenSearch, AWS Lambda) or GCP (Vertex AI, Vertex AI Agent Builder, Cloud Run, BigQuery).
  • Cloud Infrastructure Infrastructure-as-Code (IaC): Advanced experience writing and managing cloud provisioning scripts using Terraform, AWS CloudFormation, or Google Cloud Deployment Manager to deploy auto-scaling, resilient operations environments.
  • Containerization Orchestration: Production experience managing microservices via Kubernetes (EKS/GKE) and Docker to support agentic AI workers, vector indexing engines, and automated micro-tasks.
  • Observability Cloud Telemetry: Proven capability to configure full-stack observability across cloud environments using AWS CloudWatch or GCP Cloud Logging/Monitoring to trigger automated alerts and log enrichment.
  • Advanced Automation Scripting: Strong engineering capability in Python, Go, or Shell to build autonomous cloud functions (AWS Lambda/GCP Cloud Run), self-healing infrastructure scripts, and custom ITSM connectors (Jira/ServiceNow APIs).
  • Enterprise Generative AI Stack: Production execution experience deploying RAG (Retrieval-Augmented Generation) architectures using cloud vector engines (Amazon Bedrock Knowledge Bases, GCP Vertex AI Search, Pinecone, or Qdrant) for automated incident context retrieval.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Architect - Cloud AIOps
Systems Architect - Cloud AIOps

Epam Systems • Hyderabad, Chennai District, Bengaluru

On-site
INR 4,200,000 - 6,300,000
CSI Lead
CSI Lead

Opstree Global • Gurugram District

Hybrid
INR 4,500,000 - 7,500,000
Senior AI Platform Software Engineer
Senior AI Platform Software Engineer

Insight Global • Hyderabad

On-site
INR 4,500,000 - 7,000,000
AIops Lead
AIops Lead

Pineswift Technologies • Gurugram District

On-site
INR 5,000,000 - 8,000,000
Systems Integration Specialist Advisor
Systems Integration Specialist Advisor

NTT DATA North America • Dadri

On-site
INR 2,500,000 - 4,000,000
Senior Platform Engineer
Senior Platform Engineer

EPAM Systems • Chennai District

On-site
INR 2,400,000 - 4,200,000
Senior AI Ops Architect
Senior AI Ops Architect

Top Gen AI Jobs • Chennai District

On-site
INR 9,725,000 - 9,839,000
Systems Integration Specialist
Systems Integration Specialist

NTT DATA North America • Hyderabad

On-site
INR 2,000,000 - 3,500,000
AI Technical Lead
AI Technical Lead

Generac Power Systems • Pune District

On-site
INR 2,800,000 - 4,200,000
Senior Solutions Engineer
Senior Solutions Engineer

Eigenrisk • Bengaluru

Hybrid
INR 2,400,000 - 4,200,000