Observability Architect

Infinites Hr Services Pune

Hyderabad

On-site

INR 2,000,000 - 4,000,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Infinites Hr Services Pune is seeking an experienced Observability Architect to join our platform engineering and reliability team. You will design ML models for log and telemetry data, build scalable pipelines, and help shape the observability platform for AI-powered operations.

You will collaborate with SREs and product teams to translate complex challenges into scalable solutions, mentor junior data scientists, and contribute to platform standards and tooling across the stack.

Qualifications

  • 3+ years of experience in data science, ML engineering, or observability analytics.
  • Strong Python skills; production code integration in Java/Go/.NET is a plus.
  • Experience with ML frameworks: TensorFlow, PyTorch, scikit-learn.
  • Knowledge of NLP for log analysis and time-series anomaly detection.
  • Familiarity with Spark, Flink, or Kafka; graph technologies welcomed.
  • Experience with telemetry platforms like ELK, Datadog, Splunk, Prometheus.
  • Ability to communicate complex models to diverse stakeholders.
  • Exposure to AI-powered operational intelligence and agentic AI concepts.

Responsibilities

  • Design and build ML models for error and anomaly detection from logs.
  • Develop log correlation across distributed systems.
  • Build knowledge graphs of system dependencies and failures.
  • Create automated incident analysis pipelines in real time.
  • Design scalable telemetry ingestion, aggregation, and analytics pipelines.
  • Contribute to platform standards, tech choices, and engineering practices.
  • Collaborate with SRE and product teams on scalable AI-powered solutions.
  • Participate in design reviews to improve reliability, security, and performance.
  • Mentor junior data scientists in applied ML and observability.
  • Document methodologies and create dashboards for performance.
  • Stay updated on observability tech and AI workflows.

Skills

Python
ML model development
NLP for logs
Time-series anomaly detection
Graph neural networks
Distributed computing
Knowledge graphs
Telemetry platforms
Observability concepts
AI workflows
Stakeholder communication

Education

Bachelor's in Data Science or related
Master's in ML/CS (preferred)

Tools

Docker
Kubernetes
CI/CD pipelines

Job description

Position Overview

We are seeking an experienced Observability Architect to join our platform engineering and reliability team focused on developing intelligent systems for AIOps and Data Science. In addition to building machine learning models for log and telemetry data, this role will contribute to product strategy, collaborate closely with engineering teams, and help shape the architecture of our observability and operational intelligence platform.

Key Responsibilities
  • Design and build machine learning models for error detection, anomaly detection, and root cause analysis from system logs and metrics.
  • Develop log correlation algorithms identifying relationships between disparate log entries across distributed systems.
  • Build and maintain knowledge graphs representing system dependencies, service relationships, and failure patterns.
  • Create automated incident analysis pipelines ingesting raw logs, correlating events, and suggesting root causes in real time.
  • Collaborate with platform engineers to design scalable, resilient telemetry ingestion, aggregation, and analytics pipelines supporting real-time operational intelligence.
  • Contribute to defining platform standards, technology selections, and engineering practices aligned with the long-term product vision.
  • Partner with SRE, engineering and product teams to translate enterprise observability challenges into scalable, AI-powered platform solutions.
  • Participate in design reviews and help optimize platform reliability, performance, security, automation, and scalability.
  • Mentor and guide junior data scientists and cross-functional teams in applied ML and observability techniques.
  • Document model methodologies and assumptions and create dashboards for model performance and stakeholder visibility.
  • Stay abreast of emerging observability technologies, AI workflows and telemetry-driven operational intelligence trends.
  • Optional: Support the development of backend services or APIs in languages such as Python, Java, or .NET to integrate ML components with platform services.
Required Skills & Qualifications
  • Bachelor's degree in Data Science, Computer Science, Statistics, or a related field, or equivalent experience.
  • 3+ years of professional experience in data science, machine learning engineering, or observability analytics.
  • Strong programming skills in Python (preferred), with optional experience in Java, Go, or .NET for production code integration.
  • Experience developing ML models using frameworks such as TensorFlow, PyTorch, or scikit-learn.
  • Expertise in NLP techniques for log analysis and time-series anomaly detection.
  • Familiarity with distributed computing environments such as Spark, Flink, or Kafka.
  • Knowledge of knowledge graph technologies and graph neural networks, such as Neo4j, PyG, or DGL.
  • Experience with telemetry and observability platforms such as ELK, Datadog, Splunk, Prometheus, or similar tools.
  • Understanding of observability concepts, including metrics, logs, traces, dashboards, alerting, SLOs/SLIs, incident management and operational analytics.
  • Exposure to AI-powered operational intelligence, including agentic AI workflows, LLM-based assistants, graph-based dependency mapping, automated incident detection and triage, root cause analysis, and remediation recommendations.
  • Strong problem-solving skills and the ability to communicate complex models to technical and non-technical stakeholders.
Preferred Qualifications
  • Master's degree in Machine Learning, Computer Science, or a related discipline.
  • Experience in AIOps, Site Reliability Engineering, or IT Operations domains.
  • Published research or contributions to open-source ML or observability projects.
  • Knowledge of causal inference techniques for root cause analysis.
  • Familiarity with containerization technologies, including Docker and Kubernetes, and CI/CD pipelines.
  • Experience with incident management systems and on-call tooling.
  • Background working with microservices architectures or cloud platforms; Azure experience is preferred.
  • Awareness of SRE practices, self-healing automation, capacity prediction, and operational decision intelligence.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Observability Architect AIOps Data Science
Observability Architect AIOps Data Science

Mployee.me • Hyderabad

On-site
INR 800,000 - 1,300,000
Observability Architect AIOps Data Science
Observability Architect AIOps Data Science

The Glove Company • Hyderabad

Hybrid
INR 2,400,000 - 3,600,000
Observability & Sr. Observability Platform Engineer
Observability & Sr. Observability Platform Engineer

American Express Global Business Travel • Bengaluru

On-site
INR 3,500,000 - 7,500,000
AI/ML Engineer - AI Observability (ML Ops Engineer)
AI/ML Engineer - AI Observability (ML Ops Engineer)

Vconstruct • Pune District, Nagpur District

On-site
INR 2,400,000 - 4,200,000
Consultant - AI & DevOps
Consultant - AI & DevOps

SI2 Technologies Pvt Ltd • Vadodara

On-site
INR 1,500,000 - 2,100,000
DevOps Engineer
DevOps Engineer

KnowledgeWorks Global Ltd. • Mumbai

On-site
INR 1,800,000 - 3,000,000
Enterprise Observability Platform Engineer
Enterprise Observability Platform Engineer

Be a Catalyst • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Software Development Engineer
Software Development Engineer

Ciroos • Gurugram District

On-site
Confidential
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Deputy Director - AI Solution and Platforms
Deputy Director - AI Solution and Platforms

PepsiCo • Hyderabad

On-site
INR 1,000,000 - 1,500,000