Site Reliability Engineer (SRE) / Observability Engineer

N Human Resources & Management Systems

Hyderabad

Hybrid

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work
Certification reimbursement
Structured learning

Job summary

N Human Resources & Management Systems is hiring a Site Reliability Engineer with a strong Observability focus in Hyderabad. The role emphasizes building a robust monitoring stack, SLO/SLA governance, and proactive reliability improvements across distributed services.

The ideal candidate will lead instrumentation with tracing and metrics, automate toil, and drive sustainable on-call practices while partnering with global teams. Hybrid work is available from HITEC City, Hyderabad.

Qualifications

  • 8+ years in SRE/DevOps or platform engineering.
  • Hands-on Grafana dashboards, alerts, data sources.
  • Proficiency with Prometheus and PromQL.
  • Experience with log aggregation (Loki/ELK).
  • Understanding of distributed systems and containers (Kubernetes).
  • Automation with Python/Go/Bash for tooling.

Responsibilities

  • Define and report on SLOs, SLIs, and error budgets.
  • Build and optimize observability stack using Grafana/Prometheus/Loki.
  • Create dashboards and alert rules for actionable insights.
  • Lead blameless post-incident reviews to drive reliability.
  • Instrument apps with distributed tracing and metrics.
  • Automate runbooks and self-healing workflows.
  • Define on-call practices and escalation policies.
  • Evaluate new observability tools as stack evolves.

Skills

SRE experience
Grafana dashboards
PromQL
Distributed systems
Python/Go/Bash
Kubernetes
Observability strategy
Incident response

Tools

Grafana
Prometheus
Loki
Tempo
OpenTelemetry
Jaeger
VictoriaMetrics
ELK stack
Kubernetes

Job description

N Human Resources & Management Systems | Full time

Site Reliability Engineer (SRE) / Observability Engineer

We are hiring on behalf of a well-established global IT consulting and implementation firm with offices across North America, Europe, and India (HITEC City, Hyderabad). The organisation delivers technology solutions across Cloud, DevOps, SAP, and AI for enterprise clients globally and has a strong people-first, learning-oriented culture.

Role overview

We are looking for a Site Reliability Engineer with a strong Observability specialisation to drive service reliability, reduce operational toil, and build best-in-class monitoring and alerting infrastructure. The ideal candidate brings deep Grafana expertise and will take ownership of SLO/SLA definition, distributed system visibility, and driving the shift from reactive to proactive operations.

Key responsibilities
  • Define, track, and report on Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets across platform services
  • Build, maintain, and optimise observability infrastructure using Grafana, Prometheus, Loki, Tempo, and related open-source tooling
  • Develop dashboards and alerting rules that provide actionable, low-noise insights for engineering and operations teams
  • Lead blameless post-incident reviews (PIRs) and drive systemic reliability improvements from learnings
  • Partner with engineering teams to instrument applications with distributed tracing, structured logging, and custom metrics
  • Reduce operational toil through automation — scripting runbooks, auto-remediation workflows, and self-healing infrastructure
  • Define on-call practices, escalation policies, and runbooks; contribute to a sustainable on-call culture
  • Evaluate and implement new observability tooling as the stack evolves (e.g., OpenTelemetry, Jaeger, VictoriaMetrics)
Required skills & experience
  • 8+ years of combined SRE / DevOps / Platform Engineering experience
  • Strong hands-on expertise with Grafana — dashboards, alerting, data sources
  • Proficiency in Prometheus — PromQL, exporters, alertmanager
  • Experience with log aggregation using Loki, ELK stack, or equivalent
  • Solid understanding of distributed systems principles, microservices architecture, and container orchestration (Kubernetes)
  • Proficiency in Python, Go, or Bash for automation and tooling
  • Strong analytical thinking for root cause analysis and capacity planning
Good to have
  • Hands-on experience with OpenTelemetry instrumentation
  • Exposure to Grafana OnCall, Grafana Incident, or PagerDuty for incident management
  • Familiarity with eBPF-based observability tools (Cilium, Parca)
  • Azure or AWS certifications
What's on offer
  • End-to-end ownership of observability — not just maintaining dashboards
  • Hybrid work flexibility from HITEC City, Hyderabad
  • Exposure to global-scale distributed systems for international clients
  • Certification reimbursement and structured learning pathways

Experience: 8+ years

Employment type: Full-time

Specialisation: Observability – Grafana, Prometheus, Loki stack

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Engineer (SRE)
Observability Engineer (SRE)

Cloudstepin • Hyderabad

On-site
INR 1,200,000 - 1,800,000
SRE - Site Reliability Engineering
SRE - Site Reliability Engineering

Build & Hire • Pune District

On-site
INR 1,500,000 - 2,300,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Chennai District

On-site
INR 3,500,000 - 7,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Hyderabad

On-site
INR 4,200,000 - 7,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Mumbai

On-site
INR 4,000,000 - 6,000,000
Observability / SRE Engineer
Observability / SRE Engineer

Synapse Business Systems • Hyderabad, Bengaluru

On-site
INR 1,500,000 - 3,000,000
Senior SRE
Senior SRE

TecQubes Technologies • Bengaluru

Hybrid
INR 2,500,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Site Reliability Engineer
Site Reliability Engineer

Persistent • Pune District

Hybrid
INR 1,200,000 - 2,000,000
Competitive salary and benefits package
Quarterly growth opportunities
Company-sponsored higher education
+3