Site Reliability Engineer

Metlife

Hyderabad

Hybrid

INR 1,500,000 - 2,600,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Metlife Hyderabad is seeking a Site Reliability Engineer to ensure highly available, scalable data platforms and services. The role focuses on observability, automation, incident response, and collaboration with engineering, data, cloud, and operations teams to deliver reliable services.

Experience in production support and DevOps practices is essential. The candidate will design and operate large-scale data systems, develop automation scripts, and contribute to runbooks, postmortems, and

Qualifications

  • Degree in computer science, engineering, or equivalent practical experience.
  • Hands-on scripting for automation and reliability in cloud data environments.
  • Exposure to Azure data services and modern DevOps practices.

Responsibilities

  • Ensure availability, performance, scalability, and reliability of data platforms and services.
  • Design and operate large-scale data systems with automation to reduce toil.
  • Develop scripts and tools to improve monitoring, incident response, and observability.
  • Build and improve monitoring dashboards and signals for faster detection and recovery.
  • Participate in incident response, root cause analysis, and postmortems.

Skills

Python scripting
PowerShell
Bash
Spark
Automation
Observability

Education

Bachelor's degree in CS/Engineering

Tools

Azure
Azure Data Lake
Azure Synapse
Databricks
Docker
Kubernetes
Azure DevOps

Job description

Site Reliability Engineer
Role Overview

A Site Reliability Engineer (SRE) will be responsible for ensuring highly available, scalable, performant, and reliable data platforms, pipelines, and services. The role will drive operational excellence, strengthen observability, improve incident response, reduce toil through automation, and collaborate with engineering, data, cloud, and operations teams to deliver resilient services and measurable business outcomes.

Key Responsibilities
  • Ensure availability, performance, scalability, and reliability of data platforms, pipelines, and services through proactive monitoring, troubleshooting, and timely issue resolution.
  • Design, implement, and operate large-scale data systems in partnership with engineering teams, using automation and tooling to streamline operations.
  • Develop scripts, utilities, and reusable automation to reduce manual effort, minimize errors, improve efficiency, and support repeatable platform operations.
  • Build and continuously improve monitoring, alerting, dashboards, and observability signals for early detection, faster response, and improved platform health.
  • Participate in incident response, service restoration, root cause analysis, postmortems, and corrective actions to strengthen resilience and prevent recurring issues.
  • Maintain clear documentation, runbooks, processes, and operational procedures while promoting knowledge sharing and alignment with governance, controls, and production support standards.
Candidate Qualifications
  • 4+ years of experience in production support, DevOps, infrastructure, cloud operations, data platform operations, or software engineering.
  • Experience supporting business-critical systems and working within incident, problem, and change management processes.
  • Strong scripting and automation skills using Python, PowerShell, Bash, Spark, or equivalent technologies.
  • Bachelor's degree in computer science, engineering, or equivalent practical experience.
  • Exposure to hybrid cloud platforms, Azure-hosted services, regulated enterprise environments, insurance, banking, or financial services is preferred.
  • Business proficiency in English and Japanese language skills is preferred.
Tech Stack
  • Programming & Automation: Python, Spark, Bash, and PowerShell for automation, diagnostics, data processing, and operational tooling.
  • Azure Data & Cloud Platforms: Azure Data Lake, Data Factory, Synapse Analytics, Azure SQL, Cosmos DB, Databricks, and hybrid production support environments.
  • Observability & Reliability: Azure Monitor, Application Insights, Log Analytics, Splunk, AppDynamics, ELK, SLIs, SLOs, SLAs, incident management, RCA, and toil reduction.
  • DevOps & Cloud-Native Operations: Azure DevOps, GitHub, Docker, Kubernetes, AKS, CI/CD, deployments, scalability, and platform reliability.
  • Operations, Resilience & Collaboration: ServiceNow, ITSM, runbooks, disaster recovery, multi-region architecture, AI-assisted engineering tools such as GitHub Copilot or M365 Copilot, and cross-functional collaboration.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer (SRE) – DevOps Infrastructure
Site Reliability Engineer (SRE) – DevOps Infrastructure

PQAngels Technologies Pvt. Ltd. • Bengaluru

On-site
INR 1,200,000 - 2,100,000
Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

New Era Technology • Gurugram District

On-site
INR 1,500,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

PwC India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Associate Site Reliability Engineer II
Associate Site Reliability Engineer II

MetLife • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Bengaluru

On-site
INR 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Thane

On-site
INR 1,200,000 - 1,500,000
SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Navi Mumbai

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Snapmint • Gurugram District

On-site
INR 800,000 - 1,200,000