SRE Engineer- Data Platform

Metlife

Hyderabad

On-site

INR 2,500,000 - 4,000,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MetLife in Hyderabad seeks an experienced Site Reliability Engineer for the Data domain to ensure highly available, scalable data platforms, pipelines, and services. You will drive operational excellence, strengthen observability, and reduce toil through automation.

Collaborating with engineering, data, cloud, and operations teams, you will enable resilient data services and measurable business outcomes, manage incidents, and contribute to postmortems and runbooks.

Qualifications

  • Hands-on experience in SRE/DevOps for data platforms and pipelines.
  • Proficiency with Python, PowerShell, Bash for automation and tooling.
  • Experience with Azure Data services and Kubernetes for production workloads.
  • Strong incident management and observability skills.

Responsibilities

  • Data Platform Reliability: Ensure availability, performance, scalability of data platforms and pipelines.
  • Service Design and Implementation: Collaborate to operate large-scale data systems and tooling.
  • Automation and Scripting: Develop automation tools to streamline operations.
  • Monitoring and Alerting: Design and improve observability signals for early detection.

Skills

Python
Spark
PowerShell
Bash
Azure
Kubernetes
Monitoring
Automation
Git

Education

Bachelor's degree in Computer Science or related field

Tools

Azure DevOps
GitHub Enterprise
AKS
Terraform
ELK/Logs

Job description

We are offering an exciting opportunity to contribute to MetLife's digital and AI transformation journey. We are seeking an experienced Site Reliability Engineer (SRE) for the Data domain to ensure highly available, scalable, performant, and reliable data platforms, pipelines, and services. This role will drive operational excellence, strengthen observability, improve incident response, reduce toil through automation, and partner with engineering, data, cloud, and operations teams to enable resilient data services and measurable business outcomes.

Core Responsibilities:

  • Data Platform Reliability: Ensure the availability, performance, scalability, and reliability of data platforms, pipelines, and services through proactive monitoring, troubleshooting, and issue resolution.
  • Service Design and Implementation: Collaborate with engineering teams to design, implement, and operate large-scale data systems, including software and tooling to automate and streamline operations.
  • Automation and Scripting: Develop and maintain scripts, utilities, and automation tools to improve operational efficiency, reduce manual effort, minimize errors, and support repeatable data platform operations.
  • Monitoring and Alerting: Design, implement, and continuously improve monitoring, alerting, dashboards, and observability signals to enable early detection and timely resolution of issues.
  • Collaboration and Communication: Partner with engineering teams, product managers, data stakeholders, cloud teams, and operations teams to ensure services meet business requirements and align with company goals.
  • Incident Response and Management: Participate in incident response, service restoration, root cause analysis, postmortems, and corrective actions to improve resiliency and prevent recurrence.
  • Documentation and Knowledge Sharing: Maintain accurate documentation for systems, services, runbooks, processes, and operational procedures, and promote knowledge sharing across the team.
  • Governance and Day-to-Day Operations: Support day-to-day data platform operations while ensuring alignment with internal standards, global governance requirements, operational controls, and production support processes.

Skills & Experience:

  • Programming, Scripting, and Data Processing: Hands-on experience with Python, Spark, Bash, and PowerShell for automation, diagnostics, data processing support, operational tooling, and repeatable platform operations.
  • Azure Data Platform: Experience supporting Azure Data Lake Gen2, Azure Data Factory, Azure Synapse Analytics including Data Warehouse, Spark, and Pipelines, Azure SQL Database, Cosmos DB, and Databricks in production environments.
  • Observability and Monitoring: Proficiency with Azure Application Insights, Azure Log Analytics, Azure Monitor, Splunk, AppDynamics, and ELK to monitor availability, performance, pipeline health, application behavior, logs, metrics, traces, and alerts.
  • DevOps and Source Control: Working knowledge of Azure DevOps and GitHub Enterprise for repositories, CI/CD pipelines, release management, operational changes, branching strategies, and controlled production deployments.
  • Containers and Cloud-Native Operations: Experience with Docker, Kubernetes, and Azure Kubernetes Service (AKS) to support containerized workloads, platform reliability, deployment operations, scaling, and service health management.
  • ITSM and Operational Governance: Experience using ServiceNow or equivalent ITSM platforms for incident, problem, change, request, and knowledge management in a governed enterprise environment.
  • SRE and Production Reliability: Strong understanding of SLIs, SLOs, SLAs, error budgets, incident response, root cause analysis, postmortems, runbooks, automation, production readiness, and toil reduction practices.
  • Disaster Recovery & Resilience Engineering: Design and architect for multi-region systems to meet strict RTO (Service Availability) and RPO (Data Integrity) targets. You will drive down RTO via automation (Terraform, auto-scaling, DNS failover) and enforce RPO through robust replication and Continuous Data Protection (CDP) strategies.
  • AI-Assisted Engineering: Ability to use GitHub Copilot and Microsoft 365 Copilot responsibly to accelerate investigation, documentation, scripting, knowledge discovery, and operational productivity with appropriate human validation.
  • Collaboration and Execution: Strong communication, documentation, troubleshooting, evidence capture, escalation management, and cross-functional collaboration skills across engineering, data, cloud, product, and operations teams.
  • 8 - 12 years in production support, DevOps, infrastructure, cloud operations, or software engineering.
  • Experience supporting business-critical systems and working in incident, problem, and change management processes.
  • Ability to script and automate standard operational tasks using Python, PowerShell, Bash, or equivalent.
  • Bachelor's degree in computer science, engineering, or equivalent practical experience.
  • Exposure to regulated enterprise, insurance, banking, or financial services environments preferred.
  • Exposure to Hybrid cloud platforms including on-premises and Azure-hosted services.
  • Business proficiency in English; Japanese language skills are a plus.

Preferred Exposure:

  • Japanese language ability, including reading and writing.
  • Domain knowledge of life insurance business processes, data platforms, and operational support needs.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer I
Sr. Site Reliability Engineer I

MetLife • Hyderabad

Hybrid
INR 1,500,000 - 2,300,000
Associate Site Reliability Engineer II
Associate Site Reliability Engineer II

MetLife • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000
SRE Engineer
SRE Engineer

Prodapt Solutions Private Limited • Chennai District

On-site
INR 1,800,000 - 3,000,000
Senior SRE Engineer – AI Managed Services
Senior SRE Engineer – AI Managed Services

Jobtailor • Bengaluru

On-site
INR 1,800,000 - 2,800,000
DevOps & SRE Lead
DevOps & SRE Lead

Syngentagroup • Pune District

On-site
INR 4,000,000 - 7,000,000
Sr Software Engineer Syseng
Sr Software Engineer Syseng

LPL Financial Global Capability Center • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000
Sr Software Engineer Syseng
Sr Software Engineer Syseng

LPL Financial Global Capability Center • Telangana

On-site
INR 2,800,000 - 4,800,000
Senior Platform SRE
Senior Platform SRE

Tech Data Advanced Private Limited • India

On-site
INR 1,800,000 - 3,200,000