Senior Observability Architect (Prometheus and Grafana)

Keka Technologies Private Limited

Bengaluru

Hybrid

INR 3,600,000 - 7,200,000

Part time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Keka Technologies Private Limited is seeking a Senior Observability Architect to lead design and implementation of metrics, alerts, and monitoring across client environments. You will set architectural direction, standardize Prometheus/Grafana usage, and mentor delivery teams.

The role demands hands-on leadership, expertise in Prometheus, Grafana, and related tooling, and strong collaboration with stakeholders to mature observability capabilities at scale.

Qualifications

  • 15+ years in infrastructure, platform, or SRE with hands-on observability depth.
  • Expert-level Prometheus including PromQL, TSDB, retention and scaling limits.
  • Expert-level Grafana with dashboards, alerting, and data source integration.
  • Strong command of exporters (Node Exporter, cAdvisor, Blackbox Exporter) and instrumentation.
  • Experience with ITSM integrations, webhooks, and middleware for ticketing/alerting.
  • Knowledge of long-term storage (Mimir/Thanos/Cortex) and remote-write.
  • Kubernetes-native deployment using Helm and Prometheus Operator.
  • Solid Linux, networking and infra fundamentals.

Responsibilities

  • Own end-to-end architecture for Prometheus/Grafana observability platforms across client environments.
  • Define reference architectures and reusable patterns for multi-client deployments.
  • Design alerting and event-correlation pipelines and integration with ITSM workflows.
  • Map monitored assets to services, with CMDB integration as needed.
  • Lead storage design for scalable observability using Mimir/Thanos/Cortex.
  • Deploy and operate stacks in Kubernetes using kube-prometheus-stack and related exporters.
  • Establish instrumentation standards, dashboards, and alert hygiene across teams.
  • Provide senior escalation and act as technical owner on observability questions.
  • Mentor delivery teams to raise observability capability across the organization.

Skills

Prometheus
Grafana
PromQL
Alertmanager
Kubernetes
Thanos
Mimir
Cortex
Exporters
ITSM integration
Linux fundamentals
Prometheus Operator

Tools

Node Exporter
cAdvisor
Blackbox Exporter
Prometheus client libraries
Helm

Job description

Senior Observability Architect (Prometheus and Grafana)

15+years

Part-Time

Role Summary

We are looking for a Senior Observability Architect to lead the design, implementation, and maturity of metrics, alerting, and monitoring platforms across multiple client environments. As a senior, hands-on member of the team, this person will both set architectural direction and build the solution directly. The ideal candidate will own the observability strategy for current and future engagements, standardize how the organization designs, deploys, and operates Prometheus and Grafana stacks at scale, work directly with client stakeholders, and provide technical leadership to delivery teams.

Key Responsibilities
  • Own end-to-end solution architecture for Prometheus and Grafana based observability platforms, from greenfield design through to operational handover.
  • Define reference architectures and reusable patterns that can be applied consistently across multiple clients and deployment environments.
  • Architect alerting and event-correlation pipelines, including complex one-to-many correlation where a single infrastructure failure maps to multiple impacted downstream services or consumers.
  • Design and maintain the mapping between monitored assets and the services they support, including integration with CMDB or equivalent systems of record.
  • Integrate observability platforms with ITSM and ticketing systems using webhooks, APIs, and middleware, ensuring alerts flow reliably into escalation and incident workflows.
  • Lead long-term storage and scalability design using Grafana Mimir, Thanos, or Cortex, including remote-write ingestion and object-storage backends.
  • Deploy and operate observability stacks in Kubernetes environments, including the kube-prometheus-stack, Prometheus Operator, kube-state-metrics, and associated exporters.
  • Establish organizational standards for instrumentation, exporters, service discovery, dashboarding, and alert hygiene.
  • Provide senior technical escalation (L2 and above) and act as the authoritative technical owner on observability questions.
  • Mentor engineers and raise the observability capability of delivery teams across the organization.
Required Skills and Experience
  • 12-15 years in infrastructure, platform, or site reliability engineering, with deep hands-on specialization in Prometheus and Grafana based observability.
  • Expert-level Prometheus, including PromQL, recording and alerting rules, TSDB internals, retention, and scaling limitations.
  • Expert-level Grafana, including dashboard design, alerting, data source configuration, and its integration surface (webhooks, alert payload structure, and APIs).
  • Strong command of the exporter ecosystem, including Node Exporter, cAdvisor, Blackbox Exporter, and application instrumentation using Prometheus client libraries.
  • Alertmanager configuration, including deduplication, grouping, silencing, and routing to receivers such as email, Slack, PagerDuty, and Opsgenie.
  • Service discovery across Kubernetes and cloud providers.
  • Long-term storage architecture with Mimir, Thanos, or Cortex, including remote-write and object storage (S3, GCS, Azure Blob, or MinIO).
  • Kubernetes-native deployment and operations, including Helm and the Prometheus Operator.
  • Integration engineering with ITSM platforms, webhooks, and middleware for automated ticketing and escalation.
  • Solid Linux, networking, and general infrastructure fundamentals.
Nice to Have
  • Experience designing observability for managed-service or multi-tenant environments where infrastructure is shared across multiple downstream consumers.
  • Experience with GPU or high-performance compute infrastructure monitoring.
  • Familiarity with hardware-level telemetry, for example Redfish based exporters.
  • Experience with hybrid and distributed cloud platforms, including managed Kubernetes and on-premises or edge Kubernetes distributions.
  • Scripting and automation (Python, Go, or Bash) for tooling and integration.
  • Infrastructure-as-code experience (Terraform or similar).
  • Certifications such as Prometheus Certified Associate, Grafana certifications, or CKA.
Soft Skills
  • Ability to translate ambiguous client requirements into clear, defensible architecture.
  • Strong written communication, including capability notes, reference architectures, and integration documentation.
  • Comfort working directly with client technical teams and presenting to mixed technical and stakeholder audiences.
  • Self-directed ownership of complex technical problems from definition through resolution.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Enterprise Observability Platform Engineer
Enterprise Observability Platform Engineer

Be a Catalyst • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Infra Dev Specialist
Infra Dev Specialist

Cognizant • Gurugram District

On-site
INR 1,800,000 - 2,400,000
Associate Observability Architect | PST | Remote
Associate Observability Architect | PST | Remote

Embedded Shishya • Gopalganj

Hybrid
INR 13,263,000 - 15,935,000
RSUs
Remote-first culture
IT Operations Infrastructure Consultant [Prometheus & Grafana] - EG CloudOps
IT Operations Infrastructure Consultant [Prometheus & Grafana] - EG CloudOps

EG A/S • Mangaluru

On-site
INR 1,200,000 - 1,800,000
Personal and professional development
Targeted training courses in EG Academy
Best in industry employee benefits
Grafana Developer
Grafana Developer

Ericsson GmbH • Bengaluru

On-site
INR 1,800,000 - 3,600,000
Prometheus Engineer
Prometheus Engineer

Infosys • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Site Reliability Engineer (SRE) / Observability Engineer
Site Reliability Engineer (SRE) / Observability Engineer

N Human Resources & Management Systems • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work
Certification reimbursement
Structured learning
Senior Grafana Engineer
Senior Grafana Engineer

Capgemini • Hyderabad, Pune District, Bengaluru

Hybrid
INR 1,800,000 - 3,000,000
Open Source Observability Engineer
Open Source Observability Engineer

PradeepIT Consulting Services Pvt Ltd • Bengaluru

On-site
INR 1,200,000 - 1,500,000