Site Reliability Engineer (AZURE Devops)

Metlife

Hyderabad

Hybrid

INR 2,000,000 - 4,000,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Metlife in Hyderabad is looking for a Site Reliability Engineer (SRE) to ensure highly available, scalable data platforms and services. You will drive operational excellence, strengthen observability, and automate toil to deliver reliable services supporting business outcomes.

The role emphasizes incident response, RCA, runbooks, and cross-functional collaboration across engineering, data, cloud, and operations teams.

Qualifications

  • 4+ years of experience in production support, DevOps, infrastructure, cloud operations, data platform operations, or software engineering.
  • Experience working with incident, problem, and change management processes.
  • Strong scripting and automation skills using Python, PowerShell, Bash, Spark, or equivalent technologies.

Responsibilities

  • Ensure availability, performance, scalability, and reliability of data platforms, pipelines, and services through proactive monitoring, troubleshooting, and timely issue resolution.
  • Design, implement, and operate large-scale data systems in partnership with engineering teams, using automation and tooling to streamline operations.
  • Develop scripts, utilities, and reusable automation to reduce manual effort, minimize errors, improve efficiency, and support repeatable platform operations.
  • Build and continuously improve monitoring, alerting, dashboards, and observability signals for early detection, faster response, and improved platform health.
  • Participate in incident response, service restoration, root cause analysis, postmortems, and corrective actions to strengthen resilience and prevent recurring issues.
  • Maintain clear documentation, runbooks, processes, and operational procedures while promoting knowledge sharing and alignment with governance, controls, and production support standards.

Skills

Scripting & automation
Python
PowerShell
Bash
Spark

Education

Bachelor's degree in CS/Engineering

Tools

Azure Data Lake
Azure Data Factory
Synapse Analytics
Azure SQL
Cosmos DB
Databricks
Azure Monitor
AppDynamics
Splunk
ELK
Docker
Kubernetes
AKS
CI/CD

Job description

Site Reliability Engineer- Only HYDERABAD

Role Overview

A Site Reliability Engineer (SRE) will be responsible for ensuring highly available, scalable, performant, and reliable data platforms, pipelines, and services. The role will drive operational excellence, strengthen observability, improve incident response, reduce toil through automation, and collaborate with engineering, data, cloud, and operations teams to deliver resilient services and measurable business outcomes.

Key Responsibilities
  • Ensure availability, performance, scalability, and reliability of data platforms, pipelines, and services through proactive monitoring, troubleshooting, and timely issue resolution.
  • Design, implement, and operate large-scale data systems in partnership with engineering teams, using automation and tooling to streamline operations.
  • Develop scripts, utilities, and reusable automation to reduce manual effort, minimize errors, improve efficiency, and support repeatable platform operations.
  • Build and continuously improve monitoring, alerting, dashboards, and observability signals for early detection, faster response, and improved platform health.
  • Participate in incident response, service restoration, root cause analysis, postmortems, and corrective actions to strengthen resilience and prevent recurring issues.
  • Maintain clear documentation, runbooks, processes, and operational procedures while promoting knowledge sharing and alignment with governance, controls, and production support standards.
Candidate Qualifications
  • 4+ years of experience in production support, DevOps, infrastructure, cloud operations, data platform operations, or software engineering.
  • Experience supporting business-critical systems and working within incident, problem, and change management processes.
  • Strong scripting and automation skills using Python, PowerShell, Bash, Spark, or equivalent technologies.
  • Bachelor's degree in computer science, engineering, or equivalent practical experience.
  • Exposure to hybrid cloud platforms, Azure-hosted services, regulated enterprise environments, insurance, banking, or financial services is preferred.
  • Business proficiency in English and Japanese language skills is preferred.
Tech Stack
  • Programming & Automation: Python, Spark, Bash, and PowerShell for automation, diagnostics, data processing, and operational tooling.
  • Azure Data & Cloud Platforms: Azure Data Lake, Data Factory, Synapse Analytics, Azure SQL, Cosmos DB, Databricks, and hybrid production support environments.
  • Observability & Reliability: Azure Monitor, Application Insights, Log Analytics, Splunk, AppDynamics, ELK, SLIs, SLOs, SLAs, incident management, RCA, and toil reduction.
  • DevOps & Cloud-Native Operations: Azure DevOps, GitHub, Docker, Kubernetes, AKS, CI/CD, deployments, scalability, and platform reliability.
  • Operations, Resilience & Collaboration: ServiceNow, ITSM, runbooks, disaster recovery, multi-region architecture, AI-assisted engineering tools such as GitHub Copilot or M365 Copilot, and cross-functional collaboration.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE
SRE

Metlife • Hyderabad

Hybrid
INR 1,500,000 - 2,100,000
Assistant Manager - Azure Site Reliability Engineer
Assistant Manager - Azure Site Reliability Engineer

Promaynov Advisory Services Pvt. Ltd • Bengaluru

On-site
INR 1,400,000 - 2,100,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Embarkgcc Services • Bengaluru

On-site
INR 1,200,000 - 1,800,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Azure Site Reliability Engineer (SRE) + SQL - SaaS Operations
Azure Site Reliability Engineer (SRE) + SQL - SaaS Operations

Zensar Technologies • Pune District

On-site
INR 2,000,000 - 2,800,000
Site Reliability Engineer
Site Reliability Engineer

MishiPay • Bengaluru

On-site
INR 2,500,000 - 3,800,000
Site Reliability Engineer
Site Reliability Engineer

PwC India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Analyst II, Production Support
Analyst II, Production Support

fis • Pune District

On-site
INR 1,500,000 - 2,300,000
Site Reliability Engineer
Site Reliability Engineer

Finthrive • Gurugram District

Hybrid
INR 1,800,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities