Senior Observability Engineer - Grafana & Prometheus

Zensar Technologies

Bengaluru

On-site

INR 2,800,000 - 4,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zensar Technologies is seeking a Senior Observability Engineer to design, implement, and manage enterprise-scale monitoring and observability solutions for a large AWS cloud modernization program.

The role emphasizes hands-on experience with Grafana, Prometheus, AWS observability services, and end-to-end visibility across complex applications. You will build centralized platforms, develop dashboards, configure alerting, and drive reliability engineering in a fast-paced cloud environment.

Qualifications

  • Strong hands-on experience with Grafana and Prometheus.
  • Experience in Monitoring, Observability, Alerting, Dashboard Development, and Metrics Collection.
  • Expertise in troubleshooting enterprise applications and production environments.
  • Strong understanding of Monitoring & Alerting and SRE Concepts.
  • Experience with AWS Observability Services: CloudWatch, AWS Managed Grafana, AWS X-Ray.
  • Excellent analytical and problem-solving skills.
  • Experience supporting large-scale enterprise applications and cloud environments.

Responsibilities

  • Design and implement enterprise observability solutions using Grafana and Prometheus.
  • Configure alerting rules, thresholds, and notification mechanisms.
  • Build AWS-native observability capabilities using CloudWatch, X-Ray, AWS Managed Grafana, and OpenTelemetry.
  • Develop and maintain operational dashboards for infrastructure, applications, and business KPIs.
  • Create real-time dashboards for performance monitoring, SLA tracking, and operational health checks.

Skills

Grafana
Prometheus
Monitoring & Alerting
Dashboard Development
Metrics Collection
Troubleshooting
SRE Concepts
AWS Observability Services

Education

Bachelor's Degree in Computer Science, Information Technology, Engineering, or related discipline

Tools

CloudWatch
X-Ray
AWS Managed Grafana
OpenTelemetry

Job description

Zensar Technologies is looking for a highly skilled Senior Observability Engineer to design, implement, and manage enterprise-scale monitoring and observability solutions for a large AWS cloud modernization program. The ideal candidate will have strong hands-on expertise in Grafana, Prometheus, Monitoring & Alerting, Dashboard Development, Metrics Collection, and Troubleshooting, along with exposure to AWS observability services.

The role involves building a centralized observability platform to provide end-to-end visibility, reliability, performance monitoring, proactive alerting, and faster incident resolution across large-scale enterprise applications and integration environments.

Key Responsibilities
Observability Platform Engineering
  • Design and implement enterprise observability solutions using Grafana and Prometheus.
  • Configure and manage monitoring platforms for metrics collection, visualization, and reporting.
  • Build AWS-native observability capabilities using CloudWatch, X-Ray, AWS Managed Grafana, and OpenTelemetry.
  • Implement monitoring and instrumentation across distributed applications and services.
  • Ensure comprehensive visibility into application, infrastructure, and platform performance.
  • Develop and maintain operational dashboards for infrastructure, applications, and business KPIs.
  • Create real-time dashboards for performance monitoring, SLA tracking, and operational health checks.
  • Design monitoring frameworks to track system utilization, availability, and service reliability.
  • Provide actionable insights through visualization and performance analytics.
Alerting & Reliability Engineering
  • Configure alerting rules, thresholds, and notification mechanisms.
  • Implement proactive monitoring and intelligent alerting capabilities.
  • Support service reliability objectives by identifying and mitigating performance issues before business impact.
  • Drive observability best practices for incident prevention and operational excellence.
  • Perform root cause analysis using metrics, logs, traces, and dashboards.
  • Support production troubleshooting and incident investigations.
  • Work closely with engineering, operations, and platform teams to resolve critical issues.
  • Improve mean time to detect (MTTD) and mean time to resolve (MTTR).
Governance & Best Practices
  • Define observability standards, monitoring frameworks, and dashboard templates.
  • Establish logging, tracing, and correlation standards across systems.
  • Develop monitoring runbooks and operational documentation.
  • Drive continuous improvement initiatives across observability and reliability practices.
Required Skills
  • Strong hands-on experience with Grafana and Prometheus.
  • Experience in Monitoring, Observability, Alerting, Dashboard Development, and Metrics Collection.
  • Expertise in troubleshooting enterprise applications and production environments.
  • Strong understanding of:
  • Monitoring & Alerting
  • Site Reliability Engineering (SRE) Concepts
  • Experience with AWS Observability Services:
  • CloudWatch
  • AWS Managed Grafana
  • AWS X-Ray
  • Excellent analytical and problem-solving skills.
  • Experience supporting large-scale enterprise applications and cloud environments.
Good to Have
  • OpenTelemetry (OTEL) instrumentation.
  • SRE practices and reliability engineering.
  • Experience with enterprise integration platforms and cloud modernization initiatives.
  • AI-driven observability, anomaly detection, and predictive monitoring.
  • Exposure to high-volume transaction processing environments.
Preferred Qualifications
  • Bachelor's Degree in Computer Science, Information Technology, Engineering, or related discipline.
  • 8-14 years of overall IT experience.
  • Proven experience in Observability Engineering, Site Reliability Engineering (SRE), Monitoring Engineering, or Platform Operations.
  • Experience working in AWS cloud environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability Engineer - Grafana & Prometheus
Senior Observability Engineer - Grafana & Prometheus

Zensar • Pune District, Bengaluru

Hybrid
INR 1,800,000 - 3,200,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
AWS Observability Engineer
AWS Observability Engineer

Tata Consultancy Services • New Delhi, Dadri

On-site
INR 1,500,000 - 2,500,000
Enterprise Observability Platform Engineer
Enterprise Observability Platform Engineer

Be a Catalyst • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Site Reliability Engineer (SRE) / Observability Engineer
Site Reliability Engineer (SRE) / Observability Engineer

N Human Resources & Management Systems • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work
Certification reimbursement
Structured learning
AI Observability & Monitoring Engineer
AI Observability & Monitoring Engineer

Elabs Infotech • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Grafana L3 Engineer / Administrator
Grafana L3 Engineer / Administrator

Wipro • Dadri

On-site
INR 1,800,000 - 2,400,000
AWS Site Reliability Engineer (SRE)
AWS Site Reliability Engineer (SRE)

Zensar • Hyderabad, Pune District

Hybrid
INR 1,500,000 - 2,100,000
Observability / SRE Engineer
Observability / SRE Engineer

Synapse Business Systems • Hyderabad, Bengaluru

On-site
INR 1,500,000 - 3,000,000
Grafana Developer
Grafana Developer

Ericsson GmbH • Bengaluru

On-site
INR 1,800,000 - 3,600,000