Sr. Site Reliability Engineer I

MetLife

Hyderabad

On-site

INR 1,800,000 - 2,400,000

Full time

33 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MetLife is seeking a Site Reliability Engineer (SRE) to ensure reliability, availability, and performance of critical applications and platforms. The SRE will monitor production systems, respond to incidents, improve observability, maintain runbooks, and automate operational tasks.

Collaboration with engineering, cloud, and infrastructure teams will drive service readiness and adherence to SLOs, SLIs, and error budgets.

Qualifications

  • 3–7 years in production support, DevOps, or cloud operations.
  • Experience supporting business-critical systems and working in incident, problem, and change management processes.
  • Ability to script and automate standard operational tasks using Python, PowerShell, Bash, or equivalent.
  • Bachelor’s degree in computer science, engineering, or equivalent practical experience.
  • Exposure to regulated enterprise environments preferred; English proficiency required.

Responsibilities

  • Monitoring: monitor service health, dashboards, alerts, and reliability indicators.
  • Incident Response: respond to alerts, participate in bridge calls, document actions.
  • Observability Support: build dashboards, logs queries, telemetry checks, alert validation.
  • Runbook Management: document procedures, update recovery steps, share knowledge.
  • Automation: create scripts for recurring checks and data collection.
  • Problem Follow-up: support root-cause analysis and postmortems.
  • Continuous Improvement: reduce alert noise and toil, close gaps.
  • SRE Alignment: support adoption of SLOs, SLIs, SLAs, and error budgets.
  • AI Readiness: assist AI-assisted anomaly detection and correlation tools.
  • Collaboration: work with multiple teams to align performance with business goals.

Skills

Linux
Networking fundamentals
Cloud fundamentals
Production operations
ITIL incident/change processes
Scripting (Python/Bash/PowerShell)
Git
CI/CD basics
Observability basics
Monitoring & alerting
Automation
Kubernetes & Docker
Terraform (IaC) awareness
SQL for operational diagnostics
Postmortems & RCA
AIOps readiness

Education

Bachelor’s degree in CS/Engineering

Tools

Git
CI/CD pipelines
ServiceNow
Elastic/ELK
Grafana
Prometheus
Splunk
APM
Azure Monitor
Kubernetes
Docker
Terraform (IaC)

Job description

Role Reasonability

MetLife is seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of critical applications and platforms.

The SRE Engineer will monitor production systems, respond to incidents, improve observability, maintain runbooks, and automate operational tasks. Working closely with engineering, cloud, and infrastructure teams, the role supports SRE practices, operational readiness, and service reliability through SLIs, SLOs, and Error Budget management.

Core Responsibilities
  • Monitoring: Monitor service health, dashboards, alerts, and key reliability indicators for assigned applications and platforms.
  • Incident Response: Respond to alerts, support bridge calls, gather evidence, execute runbooks, communicate status, and escalat when required.
  • Observability Support: Create and maintain dashboards, log queries, telemetry checks, alert validation, and actionable monitoring signals.
  • Runbook Management: Document operational procedures, update recovery steps, validate readiness with service owners, and support knowledge sharing.
  • Automation: Create scripts for repetitive checks, data collection, remediation, operational reporting, and toil reduction.
  • Problem Follow-up: Support root cause analysis, postmortem documentation, and closure of assigned corrective/preventive action items.
  • Continuous Improvement: Identify alert noise, toil, monitoring gaps, and preventive improvements for senior SRE review.
  • SRE Alignment: Support adoption of SLOs, SLIs, SLAs, error budgets, operational readiness reviews, and production support standards.
  • AI Readiness: Use or help improve AI-assisted tools for anomaly detection, incident correlation, root cause hints, and operational knowledge retrieval.
  • Collaboration: Work with engineering, infrastructure, cloud, and application teams to align service performance with business goals.
Skills And Experience
  • Foundations: Linux, networking fundamentals, application support, cloud fundamentals, production operations, and ITIL-style incident/change processes.
  • Scripting: Python, PowerShell, Bash, or equivalent scripting for automation, diagnostics, evidence collection, and reporting.
  • Tools: Git, CI/CD basics, ServiceNow or equivalent ticketing; exposure to Elastic/ELK, Grafana, Prometheus, Splunk, APM, and Azure Monitor preferred.
  • Cloud & Containers: Azure services, Docker, Kubernetes, and hybrid cloud operations exposure; Terraform or infrastructure-as-code awareness preferred.
  • Reliability: Basic understanding of SLIs, SLOs, SLAs, error budgets, alerting, incident response, postmortems, and operational runbooks.
  • AI / AIOps Readiness: Ability to use AI-assisted investigation, anomaly detection, and correlation tools responsibly, with strong validation of evidence.
  • Database: Hands‑on SQL skills for operational diagnostics, data validation, and service health checks.
  • Execution: Disciplined follow‑through, evidence capture, documentation, escalation hygiene, and collaboration during incidents.
  • Learning Mindset: Willingness to deepen skills in cloud, Kubernetes, observability, automation, resilience engineering, and secure operations.
Minimum Qualifications
  • 3–7 years in production support, DevOps, infrastructure, cloud operations, or software engineering.
  • Experience supporting business‑critical systems and working in incident, problem, and change management processes.
  • Ability to script and automate standard operational tasks using Python, PowerShell, Bash, or equivalent.
  • Bachelor’s degree in computer science, engineering, or equivalent practical experience.
  • Exposure to regulated enterprise, insurance, banking, or financial services environments preferred.
  • Business proficiency in English; Japanese language skills are a plus.
Preferred Exposure
  • Hybrid cloud platforms including on‑premises and Azure‑hosted services.
  • Observability platforms such as ELK/Elastic, Grafana, Prometheus, Splunk, Azure Monitor, and Azure Application Insights.
  • GitHub, Azure DevOps, pipelines, repositories, and operational change controls.
  • Kubernetes‑based production services and containerized application support.
  • SRE practices including operational readiness reviews, service health reviews, and toil reduction initiatives.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Associate Site Reliability Engineer II
Associate Site Reliability Engineer II

MetLife • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Senior Associate Site Reliability Engineer
Senior Associate Site Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Hyderabad

On-site
INR 1,400,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer (SRE) – Core IT Infrastructure
Site Reliability Engineer (SRE) – Core IT Infrastructure

TECEZE • Chennai District

On-site
INR 1,000,000 - 2,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Augusta Infotech • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000