AIOps Engineer_GCP

Horizon Industries International Limited

Gurugram District

On-site

INR 2,000,000 - 4,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Horizon Industries International Limited is seeking an experienced AIOps Engineer to design, implement, and optimize AI-driven IT operations on Google Cloud Platform. The role focuses on anomaly detection, event correlation, predictive incident management, and automated remediation for production ML workloads.

You will build end-to-end AI/ML operations pipelines, deploy models with Vertex AI, and collaborate with SREs, data scientists, and DevOps to embed automation.

Qualifications

  • 5+ years of experience in DevOps/SRE/Data Engineering with AIOps and MLOps.
  • 3+ years hands-on with GCP services (Vertex AI, BigQuery, Dataflow, Pub/Sub, GKE, Cloud Operations) in production.
  • Strong Python programming.
  • Experience with observability/monitoring tools (Cloud Monitoring, Prometheus, Grafana, ELK/Splunk, Datadog) and distributed computing frameworks.
  • Hands-on experience with CI/CD pipelines, Terraform/IaC, Kubernetes (GKE), and automation tools.
  • Exposure deploying production Generative AI use cases including prompt engineering and RAG frameworks.
  • Familiarity with Pub/Sub, Kafka or similar messaging systems.
  • Strong problem-solving, iteration, and communication skills.
  • Ability to collaborate across teams and explain AI-driven automation to diverse stakeholders.

Responsibilities

  • Design, develop, and maintain AIOps pipelines on GCP for telemetry data to enable intelligent monitoring.
  • Build and deploy ML models for anomaly detection, event correlation, and predictive incident management using Vertex AI and BigQuery ML.
  • Collaborate with data scientists, SREs, software engineers, and DevOps to embed AI-driven automation in IT operations.
  • Automate end-to-end ML and operations workflows using Vertex AI Pipelines, Kubeflow, or Cloud Composer.
  • Implement CI/CD pipelines using Cloud Build, GitHub Actions, and Terraform for automated deployment and monitoring.
  • Utilize Pub/Sub, Dataflow, and Kafka for real-time telemetry ingestion and streaming analytics.
  • Optimize GCP infrastructure (GKE, Compute Engine, BigQuery, Cloud Storage) for scalability and cost efficiency.
  • Manage GitHub repositories for version control on AIOps and ML projects.
  • Integrate with ITSM/ITOM platforms and various data connectors to access operational data.
  • Develop and maintain documentation for AIOps pipelines, infrastructure, and runbooks.
  • Stay up to date on AIOps, SRE, MLOps, and cloud engineering trends.

Skills

DevOps/SRE
Data Engineering
Problem-solving
Communication
Python

Education

Bachelor's or Master's in CS/Engineering

Tools

GCP Vertex AI
BigQuery
Dataflow
Pub/Sub
GKE
Cloud Monitoring
Cloud Logging
Cloud Trace
Terraform
Kubeflow
Cloud Composer
GitHub
Kafka
Prometheus
Grafana

Job description

We are looking for an experienced AIOps Engineer with expertise in Google Cloud Platform (GCP), AIOps/MLOps/LLMOps, intelligent observability and monitoring, CI/CD, Python, Pub/Sub/Kafka, distributed computing, GitHub, data pipelines, and GCP services such as Vertex AI, BigQuery, Dataflow, GKE, and Cloud Operations Suite (Cloud Monitoring, Cloud Logging, Cloud Trace). This role will involve designing, implementing, and optimizing AI-driven IT operations solutions - including anomaly detection, event correlation, predictive incident management, and automated remediation - and ensuring smooth, reliable operation of platforms and ML workloads in production environments on GCP.

Responsibilities
  • Design, develop, and maintain AIOps pipelines on GCP for ingesting, processing, and analyzing telemetry data (logs, metrics, traces, events) to enable intelligent monitoring and observability.
  • Build and deploy ML models for anomaly detection, event correlation, root-cause analysis, and predictive incident management using Vertex AI and BigQuery ML.
  • Collaborate with data scientists, SREs, software engineers, and DevOps teams to embed AI-driven automation into IT operations using best practices in AIOps and MLOps.
  • Automate end-to-end ML and operations workflows, including data preprocessing, model training, evaluation, deployment, and automated remediation, using tools like Vertex AI Pipelines, Kubeflow, or Cloud Composer (Apache Airflow).
  • Implement CI/CD pipelines using Cloud Build, GitHub Actions, and Terraform (IaC) for automated deployment, testing, and monitoring of models and services on GCP.
  • Utilize Pub/Sub, Dataflow, and Kafka for real-time telemetry ingestion, streaming analytics, and event-driven automation.
  • Optimize GCP infrastructure (GKE, Compute Engine, BigQuery, Cloud Storage) for scalability, performance, reliability, and cost efficiency.
  • Manage GitHub repositories for version control and collaboration on AIOps and machine learning projects.
  • Integrate with ITSM/ITOM platforms (e.g., ServiceNow) and various data connectors to access and process operational data from different sources.
  • Develop and maintain documentation for AIOps pipelines, GCP infrastructure, runbooks, and processes.
  • Stay up to date on emerging technologies and best practices in AIOps, SRE, machine learning operations, and cloud engineering.
Qualifications
  • 5+ Years of prior experience in DevOps/SRE/Data Engineering, with strong exposure to AIOps and MLOps.
  • 3+ Years of hands‑on experience with GCP services (Vertex AI, BigQuery, Dataflow, Pub/Sub, GKE, Cloud Operations Suite) in production environments.
  • Strong proficiency in Python programming language.
  • Experience with observability and monitoring tools (Cloud Monitoring, Prometheus, Grafana, ELK/Splunk, Datadog, or similar) and distributed computing frameworks.
  • Hands‑on experience with CI/CD pipelines, Terraform/IaC, Kubernetes (GKE), and automation tools.
  • Exposure in deploying a use case in production leveraging Generative AI involving prompt engineering and RAG Framework (e.g., LLM-assisted incident summarization or intelligent ticket triage).
  • Familiarity with Pub/Sub, Kafka, or similar messaging systems.
  • Strong problem‑solving skills and the ability to iterate and experiment to optimize AI model behavior and operational automation.
  • Excellent problem‑solving skills and attention to detail.
  • Ability to communicate effectively with diverse clients/stakeholders.
Education Background
  • Bachelor’s or master’s degree in computer science, Engineering, or a related field.

Folks with shorter notice period to be preferred.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior MLE - MLOps, Python, GCP, VertexAI, GKE
Senior MLE - MLOps, Python, GCP, VertexAI, GKE

UPS • Chennai District

On-site
INR 800,000 - 1,200,000
AI Engineer
AI Engineer

Tech Mahindra • Pune District

On-site
INR 1,800,000 - 2,400,000
Lead GCP MLOps Engineer
Lead GCP MLOps Engineer

dentsu • Maharashtra

On-site
INR 2,800,000 - 5,200,000
DevOps SE II - GCP & AI
DevOps SE II - GCP & AI

Keywords Studios • Maharashtra

On-site
INR 1,500,000 - 2,000,000
Senior - MLOPS Engineer
Senior - MLOPS Engineer

HCA Healthcare - India • Hyderabad

On-site
INR 1,500,000 - 2,000,000
Lead ML Devops Engineer
Lead ML Devops Engineer

Dentsu Global Services • Maharashtra

On-site
INR 3,000,000 - 4,000,000
Senior MLE - MLOps Python GCP VertexAI GKE
Senior MLE - MLOps Python GCP VertexAI GKE

UPS • Thiruvallur District

On-site
INR 1,800,000 - 3,000,000
GCP Devops Engineer
GCP Devops Engineer

Capgemini • Pune District, Delhi, Dadri

Hybrid
INR 1,200,000 - 2,400,000
AIOps Engineer
AIOps Engineer

Cloudly Inc • India

On-site
INR 1,000,000 - 1,500,000
Two annual festive bonuses
Health insurance
Fully subsidized lunch and snacks
GCP Infrastructure Engineer - Google Cloud Terraform Python Bash GKE CI CD
GCP Infrastructure Engineer - Google Cloud Terraform Python Bash GKE CI CD

UPS • Thiruvallur District

On-site
INR 4,000,000 - 7,000,000