Observability Engineer

Provate Inc.

Hyderabad

On-site

INR 1,000,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology solutions provider is seeking a Reliability Engineer in Hyderabad to design and enhance observability for applications. The role includes building systems, collaborating with teams to improve performance, and automating SRE practices. Candidates should possess programming skills in languages such as Go and have experience with AWS, Docker, and observability tools like Grafana. Excellent analytical skills and communication abilities are essential for influencing monitoring standards across teams.

Qualifications

  • Able to contribute tools and advise on application-level instrumentation improvements.
  • Experience with observability tools for collecting and analyzing telemetry data.
  • Strong analytical skills for interpreting data and identifying trends.

Responsibilities

  • Design and maintain infrastructure for observability systems.
  • Develop dashboards, alerts, and visualizations for system performance.
  • Collaborate with teams to define SLIs/SLOs.
  • Advocate for consistent instrumentation and effective alerting practices.
  • Continuously assess performance of observability tools.
  • Engineer automation capabilities for SRE best practices.
  • Support incident response and contribute to post-incident reviews.

Skills

Programming experience in Go, Python, Java, or Node.js
Observability tooling expertise with Loki, Grafana, Tempo, etc.
Cloud experience with AWS
Familiarity with Docker and Kubernetes
Experience with Terraform, Ansible, Chef, or SCCM
Strong understanding of Linux and shell scripting
Deep understanding of tracing technology
Knowledge of SLIs, SLOs, and Error Budgets

Job description

We’re looking for a Reliability Engineer to help design, build, and scale our observability platform and drive improvements through our SRE program. In this hands‑on role, you’ll contribute directly to the systems that monitor the health and performance of our applications, helping engineers understand system behavior, troubleshoot issues quickly, and continuously improve service reliability.

You’ll be instrumental in shaping how we observe, measure, and improve our systems, helping define a culture of accountability, ownership, and operational excellence. We’re a team that values curiosity, collaboration, and a growth mindset. This is a great opportunity to influence both the tooling and practices that support large‑scale, reliable software delivery.

Another key deliverable for this role is to help us improve reliability within our critical applications. The reliability engineer will collaborate with application teams to determine what needs to happen to deliver high levels of uptime and performance for our customers.

Key Responsibilities:
  • Build and scale observability systems: Design and maintain infrastructure for collecting, aggregating, and analyzing telemetry data (metrics, logs, and traces).
  • Enable actionable insights: Develop dashboards, alerts, and visualizations that turn raw data into clear, meaningful information for engineers, SREs, and business stakeholders.
  • Collaborate across teams: Partner with engineering, operations, and SRE teams to define SLIs/SLOs and improve visibility into system performance and health.
  • Drive best practices: Advocate for and support consistent instrumentation, effective alerting, and strong observability practices across engineering teams.
  • Optimize systems and tools: Continuously assess performance, usage, and cost of observability tools, identifying opportunities for improvement and efficiency.
  • Automate: Engineer capabilities that will drive the adoption of SRE principles and best practices into what is deployed within the environment.
  • Improve: In collaboration with engineering teams develop plans to improve the reliability of applications and infrastructure and assist these teams with the engineering of these improvements.
  • Support incident response: Participate and help improve the incident response process, reducing MTTR and contributing to post‑incident reviews and root cause analysis.

Required Skills & Experience:

Technical Skills

  • Programming experience in languages like Go, Python, Java, or Node.js. Able to contribute tools and advise on application‑level instrumentation improvements.
  • Observability tooling expertise with these tools: Loki, Grafana, Tempo, Mimr, Cloudwatch, ClickStack, VictoriaMetrics, Groundcover, Libre.
  • Cloud experience with AWS and services like EC2, EKS, ECS, VPC networking.
  • Containers & orchestration: Familiarity with Docker and Kubernetes.
  • Infrastructure as Code & automation: Experience with tools like Terraform, Ansible, Chef, or SCCM to manage observability infrastructure at scale.
  • Linux systems knowledge: Strong understanding of Linux, shell scripting, and the storage/networking stack.
  • Tracing: Deep understanding of tracing technology and OpenTelemetry.
  • SRE Practices: SLIs, SLOs, Error Budgets, and Failure Domains.

Soft Skills

  • Strong analytical skills for interpreting data and identifying trends or anomalies.
  • Clear and effective communication—both written and verbal—for working with technical and non‑technical stakeholders.
  • Able to influence different teams on monitoring standards and how to improve the reliability of their applications.

By submitting this form you agree to Provate's Privacy Policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Chennai District

On-site
INR 3,500,000 - 7,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Hyderabad

On-site
INR 4,200,000 - 7,000,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Mumbai

On-site
INR 4,000,000 - 6,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Brillio • Bengaluru Urban

On-site
INR 1,200,000 - 2,000,000
Enterprise Observability Platform Engineer
Enterprise Observability Platform Engineer

Be a Catalyst • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Senior Platform SRE
Senior Platform SRE

IG Infotech • Bengaluru

On-site
INR 1,200,000 - 1,600,000
Observability / SRE Engineer
Observability / SRE Engineer

Synapse Business Systems • Hyderabad, Bengaluru

On-site
INR 1,500,000 - 3,000,000