AIOps Engineer

Cloudly Inc

India

On-site

INR 1,000,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Two annual festive bonuses
Health insurance
Fully subsidized lunch and snacks

Job summary

Cloudly Inc in India is looking for an AIOps Engineer to enhance its AI products through operational efficiency in infrastructure. The candidate will design AIOps capabilities, develop ML models for predictive analysis, and collaborate directly with various teams.

This role requires a strong background in ML, Python programming, and observability tools, all contributing to faster incident detection and operational insight.

Cloudly Inc offers competitive salary, generous leave, health insurance, and collaborative opportunities with US teams.

Qualifications

  • 3 to 5 years of experience in ML or data engineering with production operations.
  • Strong proficiency in Python for building ML models.
  • Deep familiarity with observability tools like Prometheus and Grafana.

Responsibilities

  • Design and build AIOps capabilities for anomaly detection and alerting.
  • Develop ML models for predictive failure and performance monitoring.
  • Integrate AIOps with existing incident management workflows.

Skills

Machine Learning
Python
Observability tooling
Incident management

Education

Bachelor's or Master's degree in Computer Science, Engineering, Data Science, or related field

Tools

Prometheus
Grafana
Kubernetes
Kafka

Job description

As AIOps Engineer at CloudlyIO, you sit at the intersection of artificial intelligence and operational excellence. You will build and operate the intelligence layer on top of our infrastructure and application observability stack, applying machine learning and automation to transform raw telemetry into actionable operational insight, faster incident detection, predictive alerting, and autonomous remediation where appropriate.This role is central to how CloudlyIO operates its own AI products and also directly informs the capabilities we build into CloudlyMELT for our customers. You will both use AIOps tooling and help shape what great AIOps looks like at production scale.This is a role for someone who understands both ML techniques and production operations deeply, and who finds it genuinely exciting to apply one to improve the other.

Job Requirement

AIOps Platform Development
  • Design and build AIOps capabilities across CloudlyIO's observability and operations stack, including anomaly detection, intelligent alerting, event correlation, and root cause analysis automation
  • Develop ML models for predictive failure detection, capacity forecasting, and performance degradation early warning across cloud infrastructure and AI workloads
  • Build and maintain data pipelines that ingest telemetry from infrastructure, applications, and AI systems into unified operational intelligence systems
  • Implement LLM-powered operations tooling including natural language incident summarization, automated runbook suggestion, and root cause explanation
  • Evolve CloudlyIO's monitoring posture from reactive alerting to proactive and predictive operations using ML-driven signal analysis
  • Reduce alert noise through intelligent event correlation, deduplication, and suppression
  • Build dashboards and operational intelligence interfaces that give engineering teams clear, actionable visibility into system health and predicted risks
  • Build and maintain automated incident triage and enrichment pipelines that accelerate mean time to detection (MTTD) and mean time to resolution (MTTR)
  • Develop post-incident analysis tooling that identifies patterns across historical incidents and surfaces systemic improvement opportunities
  • Integrate AIOps capabilities with existing incident management workflows and on-call tooling
  • Work closely with CloudOps, DevOps, SecOps, and MLOps teams to embed AI-driven intelligence throughout the operational stack
  • Collaborate with the CloudlyMELT product team to ensure internal AIOps practices inform and improve our customer-facing observability product
  • Evaluate and recommend AIOps tools, frameworks, and approaches as the discipline and our needs evolve
  • Document all AIOps systems, models, and operational procedures clearly and maintain them as systems change
YOU MAY BE A GOOD FIT IF YOU HAVE
  • 3 to 5 years of experience combining ML or data engineering with production operations, platform engineering, or site reliability engineering
  • Strong proficiency in Python with hands‑on experience building ML models for anomaly detection, time series forecasting, or classification in operational contexts
  • Deep familiarity with observability tooling including Prometheus, Grafana, Elasticsearch, Kibana, CloudWatch, and Datadog
  • Experience working with large‑scale telemetry data including metrics, logs, and traces
  • Working knowledge of cloud infrastructure on AWS, with familiarity in how infrastructure components fail and how those failures manifest in telemetry
  • Experience integrating LLMs or ML models into operational workflows and tooling
  • Strong understanding of incident management practices including MTTD, MTTR, and post‑incident review processes
  • Comfort working across teams and translating operational needs into ML problem definitions
PREFERRED QUALIFICATIONS
  • Experience with GPU infrastructure observability and AI workload performance monitoring
  • Familiarity with distributed tracing frameworks such as Jaeger or OpenTelemetry
  • Experience with streaming data platforms such as Kafka or Kinesis for real‑time telemetry processing
  • Knowledge of Kubernetes operational patterns and failure modes
  • Experience contributing to or building commercial AIOps or observability products
  • Familiarity with chaos engineering and resilience testing as inputs to operational intelligence
  • Bachelor's or Master's degree in Computer Science, Engineering, Data Science, or a related field
COMPENSATION & BENEFITS
  • Salary: Competitive and negotiable based on experience
  • Two annual festive bonuses, each equivalent to half a month's salary
  • Two‑day weekends, 10 days casual leave, 10 days sick leave, and 14 public holidays per CloudlyIO's global holiday calendar
  • Fully subsidized lunch and evening snacks, plus tea and coffee throughout the day
  • Health insurance
  • Direct collaboration with US clients and teams, working on real enterprise AI infrastructure from day one
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CloudOps Engineer
CloudOps Engineer

Cloudly Inc • India

Hybrid
INR 800,000 - 1,200,000
Two annual festive bonuses
Fully subsidized lunch and snacks
Health insurance
Product Manager, CloudlyMELT
Product Manager, CloudlyMELT

Cloudly Inc • India

On-site
INR 5,725,000 - 7,634,000
Performance-based commission
Two annual festive bonuses
Fully subsidized lunch and snacks
DevOps Engineer
DevOps Engineer

Cloudly Inc • India

On-site
INR 800,000 - 1,200,000
Health insurance
Two annual bonuses
Fully subsidized lunch
+1
Sales Area Lead, CloudlyMELT
Sales Area Lead, CloudlyMELT

Cloudly Inc • India

On-site
INR 1,500,000 - 2,500,000
Competitive salary
Performance-based commission
Annual bonuses
+2
Solution Engineer, CloudlyMELT
Solution Engineer, CloudlyMELT

Cloudly Inc • India

On-site
INR 1,200,000 - 2,000,000
Performance-based commission
Two annual festive bonuses
Fully subsidized lunch
+1
Cloud Solution Architect
Cloud Solution Architect

Cloudly Inc • India

On-site
INR 6,679,000 - 10,497,000
Competitive salary
Two annual bonuses
Health insurance
+1
AIOps Engineer_GCP
AIOps Engineer_GCP

Horizon Industries International Limited • Gurugram District

On-site
INR 2,000,000 - 4,200,000
ML Engineer, CloudlyMELT
ML Engineer, CloudlyMELT

Cloudly Inc • India

On-site
INR 100,000 - 150,000
Competitive base salary
Performance-based commission structure
Two annual festive bonuses
+1
Sales Executive, CloudlyMELT
Sales Executive, CloudlyMELT

Cloudly Inc • India

On-site
INR 700,000 - 1,200,000
Competitive base salary
Performance-based commission
Two annual bonuses
+2
ML Engineer, CloudlyPulse
ML Engineer, CloudlyPulse

Cloudly Inc • India

On-site
INR 1,200,000 - 1,600,000
Competitive salary, negotiable based on experience
Performance-based commission structure
Two annual festive bonuses
+2