Observability Architect AIOps Data Science

Mployee.me

Hyderabad

On-site

INR 800,000 - 1,300,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mployee.me is seeking an experienced Observability Architect to join the Platform Engineering and Reliability team in Hyderabad. You will work on AI-powered observability for AIOps, data science, and operational intelligence, building scalable telemetry and RCA capabilities.

You will collaborate with SRE, Platform Engineering, Data Science, Product, and Software Engineering to define architecture, standards, and best practices across telemetry ingestion, analytics, and automation.

Qualifications

  • Bachelor's degree in Computer Science, Data Science, Statistics, or related field.
  • 3+ years of professional experience in Data Science, ML Engineering, or Observability Analytics.
  • Strong Python programming experience.

Responsibilities

  • Design and develop ML models for anomaly detection, error detection, and root cause analysis from logs, metrics, and telemetry.
  • Develop log correlation and event analysis algorithms across distributed systems.
  • Build and maintain knowledge graphs representing dependencies, services, and failure patterns.

Skills

Python
ML/AI observability
NLP
Time-series analytics
Knowledge Graphs
Graph Neural Networks
Neo4j
PyG
DGL
Elastic Stack
Datadog
Splunk
Prometheus
Kafka
Spark
Flink
Kubernetes
Docker
Azure
CI/CD
RCA
SRE practices
AIOps

Education

Bachelor's degree in Computer Science/Data Science/Statistics
Master's degree (preferred)

Tools

Neo4j
PyG
DGL
Elastic Stack
Datadog
Splunk
Prometheus
Kafka
Spark
Flink

Job description

Observability Architect AIOps & Data Science

Location: [Location]
Experience: 3+ Years
Employment Type: Full-Time
Work Mode: [On-site/Hybrid/Remote]

Position Overview

We are looking for an experienced Observability Architect to join our Platform Engineering and Reliability team, focused on building intelligent systems for AIOps, Data Science, and Operational Intelligence.

The role combines observability, machine learning, AI, telemetry analytics, and platform architecture to develop intelligent solutions for anomaly detection, incident analysis, log correlation, root cause analysis, and automated operational insights.

You will work closely with SRE, Platform Engineering, Data Science, Product, and Software Engineering teams to design scalable observability solutions and shape the architecture and technology strategy for an AI-powered operational intelligence platform.

Key Responsibilities

  • Design and develop ML models for anomaly detection, error detection, and root cause analysis using logs, metrics, and telemetry data.
  • Develop log correlation and event analysis algorithms across distributed systems.
  • Build and maintain knowledge graphs representing system dependencies, service relationships, and failure patterns.
  • Develop automated incident analysis pipelines that ingest telemetry, correlate events, and recommend potential root causes.
  • Design scalable and resilient telemetry ingestion, aggregation, and analytics pipelines for real-time operational intelligence.
  • Define platform standards, technology selections, architecture patterns, and engineering best practices.
  • Collaborate with SRE, engineering, and product teams to solve complex enterprise observability challenges using AI.
  • Contribute to architecture and design reviews focused on reliability, scalability, security, performance, and automation.
  • Mentor junior data scientists and cross-functional teams in ML, AIOps, and observability practices.
  • Document ML methodologies, assumptions, performance metrics, and model outcomes.
  • Build dashboards and reporting mechanisms for model performance and operational visibility.
  • Stay current with emerging technologies in observability, AI workflows, AIOps, and telemetry-driven operational intelligence.
  • Optionally develop backend services/APIs using Python, Java, or .NET to integrate ML capabilities with platform services.

Required Skills & Qualifications

  • Bachelor's degree in Computer Science, Data Science, Statistics, or a related field, or equivalent experience.
  • 3+ years of professional experience in Data Science, Machine Learning Engineering, or Observability Analytics.
  • Strong hands-on programming experience with Python.
  • Experience developing ML models using TensorFlow, PyTorch, or scikit-learn.
  • Strong understanding of NLP techniques for log analysis and time-series anomaly detection.
  • Experience with distributed technologies such as Apache Spark, Flink, or Kafka.
  • Knowledge of Knowledge Graphs and Graph Neural Networks (GNNs).
  • Experience with technologies such as Neo4j, PyG, or DGL is highly desirable.
  • Hands-on experience with observability platforms such as:
    • ELK / Elastic Stack
    • Datadog
    • Splunk
    • Prometheus
    • or equivalent observability platforms.
  • Strong understanding of metrics, logs, traces, dashboards, alerting, SLOs, SLIs, incident management, and operational analytics.
  • Exposure to AI-powered operational intelligence, including:
    • Agentic AI workflows
    • LLM-based operational assistants
    • Graph-based dependency mapping
    • Automated incident detection and triage
    • Root Cause Analysis (RCA)
    • Remediation recommendations
  • Strong analytical, problem-solving, and communication skills.

Preferred Qualifications

  • Master's degree in Machine Learning, Computer Science, Data Science, or a related discipline.
  • Experience in AIOps, SRE, IT Operations, or Observability Engineering.
  • Experience with causal inference techniques for root cause analysis.
  • Published research or contributions to ML, AI, AIOps, or observability open-source projects.
  • Experience with Docker, Kubernetes, and CI/CD pipelines.
  • Familiarity with incident management platforms and on-call tooling.
  • Experience working with microservices and cloud-native architectures.
  • Experience with Azure is preferred.
  • Understanding of SRE practices, self-healing automation, capacity prediction, and operational decision intelligence.

What You’ll Work On

Observability Telemetry ML/AI Correlation Knowledge Graphs RCA Incident Intelligence Automated Remediation

If you are passionate about combining Observability, Machine Learning, GenAI, and AIOps to build intelligent and resilient enterprise platforms, we would like to hear from you.

Keywords

Observability | AIOps | Python | Machine Learning | NLP | Anomaly Detection | Root Cause Analysis | Log Analytics | Time-Series | Knowledge Graph | GNN | Neo4j | PyTorch | TensorFlow | Kafka | Spark | Flink | ELK | Splunk | Datadog | Prometheus | LLM | Agentic AI | SRE | Kubernetes | Azure

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Observability Architect AIOps Data Science
Observability Architect AIOps Data Science

The Glove Company • Hyderabad

Hybrid
INR 2,400,000 - 3,600,000
Observability Architect
Observability Architect

Infinites Hr Services Pune • Hyderabad

On-site
INR 2,000,000 - 4,000,000
Consultant - AI & DevOps
Consultant - AI & DevOps

SI2 Technologies Pvt Ltd • Vadodara

On-site
INR 1,500,000 - 2,100,000
AI Engineer
AI Engineer

Allegis Group • Hyderabad, Bengaluru

Hybrid
INR 1,200,000 - 2,500,000
Observability & Sr. Observability Platform Engineer
Observability & Sr. Observability Platform Engineer

American Express Global Business Travel • Bengaluru

On-site
INR 3,500,000 - 7,500,000
AI/ML Engineer - AI Observability (ML Ops Engineer)
AI/ML Engineer - AI Observability (ML Ops Engineer)

Vconstruct • Pune District, Nagpur District

On-site
INR 2,400,000 - 4,200,000
Systems Architect - Cloud AIOps
Systems Architect - Cloud AIOps

EPAM Systems • Coimbatore District

On-site
INR 5,500,000 - 7,500,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Systems Architect - Cloud AIOps
Systems Architect - Cloud AIOps

Epam Systems • Hyderabad, Chennai District, Bengaluru

On-site
INR 4,200,000 - 6,300,000
Systems Architect - Cloud AIOps
Systems Architect - Cloud AIOps

EPAM Systems • Maharashtra

On-site
INR 4,200,000 - 7,000,000