Site Reliability Engineer

MetLife México

Fatih

On-site

TRY 320,000 - 540,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Private health insurance
Pension plan
Work from home allowance
Heritage Day off

Job summary

MetLife México is seeking a Site Reliability Engineer (SRE) – Observability & Elastic to join a high‑performing engineering organization that ships resilient, secure, and observable platforms across cloud and hybrid environments.

You will build and maintain observability pipelines (logs, metrics, traces) using Elastic Stack, create dashboards and alerts, define SLOs, and drive automated remediation in collaboration with development, security, and platform teams.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or related field.
  • 3+ years of experience in SRE, DevOps, or production engineering roles.
  • Strong hands‑on experience with Elastic Stack (Elasticsearch, Logstash, Kibana).
  • Proficiency in monitoring tools and observability frameworks (metrics, distributed tracing, logging).

Responsibilities

  • Design, implement, and manage end‑to‑end observability solutions (metrics, logs, traces).
  • Build and maintain Elastic Stack based logging and monitoring platforms.
  • Develop dashboards, alerts, and visualization layers for proactive issue detection.
  • Define and continuously improve SLIs, SLOs, and alerting strategies.
  • Enable log, metric, and trace correlation to improve troubleshooting efficiency.

Skills

SRE/DevOps experience
English proficiency
Python scripting
Distributed systems
CI/CD familiarity

Education

Bachelor’s degree in Computer Science, Engineering

Tools

Elastic Stack (Elasticsearch, Logstash, Kibana)
Prometheus
Grafana
Azure Monitor
App Insights
Splunk

Job description

Site Reliability Engineer (SRE) – Observability & Elastic

Site Reliability Engineer (SRE) – Observability & Elastic

The Team You Will Join

You will be part of a high-performing engineering organization responsible for delivering resilient, secure, and observable platforms. Our team works at the intersection of software engineering, infrastructure, and security, ensuring that critical systems are highly available, well‑monitored, and continuously optimized. You will collaborate closely with product teams, security experts, and platform engineers to build a strong observability and reliability culture across the organization.

The Opportunity
  • Work on enterprise‑scale, mission‑critical systems serving real business operations
  • Build and enhance observability capabilities (logs, metrics, traces) to improve system reliability and transparency
  • Utilize AI‑powered analytics on observability data (logs, metrics, traces) to detect anomalies, accelerate root cause analysis, and improve operational intelligence
  • Design and implement scalable and resilient platform solutions in cloud and hybrid environments
  • Collaborate with cross‑functional teams (development, infrastructure, security) to improve system reliability and performance
  • Contribute to automation‑first operations, reducing manual effort and increasing efficiency
  • Gain hands‑on experience with modern SRE practices including SLOs, incident management, and reliability engineering
  • Participate in building a data‑driven engineering culture using observability insights
  • Be part of a global organization with modern engineering standards, tools, and practices
Key Responsibilities
Observability & Monitoring
  • Design, implement, and manage end‑to‑end observability solutions (metrics, logs, traces)
  • Build and maintain Elastic Stack (ELK / OpenSearch) based logging and monitoring platforms
  • Develop dashboards, alerts, and visualization layers for proactive issue detection
  • Define and continuously improve SLIs, SLOs, and alerting strategies
  • Enable log, metric, and trace correlation to improve troubleshooting efficiency
Reliability Engineering
  • Ensure high availability, scalability, and performance of distributed systems
  • Drive adoption of reliability practices such as incident retrospectives and proactive monitoring
  • Participate in incident response, root cause analysis, and resilience improvement initiatives
  • Implement automated remediation and self‑healing mechanisms
Security & DevSecOps
  • Integrate security monitoring and logging (SIEM‑like use cases) into observability platforms
  • Collaborate with security teams on threat detection, anomaly monitoring, and audit logging
  • Contribute to DevSecOps practices, embedding security into CI/CD pipelines
  • Support audit readiness and compliance reporting through structured logging and monitoring
Automation & Platform Engineering
  • Automate operational workflows to reduce toil and increase efficiency
  • Contribute to the improvement of CI/CD pipelines and release processes
  • Support on‑call operations and continuously improve alert quality and signal‑to‑noise ratio
  • Develop Python scripts for synthetic monitoring and testing
Required Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or related field
  • 3+ years of experience in SRE, DevOps, or production engineering roles
  • Good command of English
  • Strong hands‑on experience with Elastic Stack (Elasticsearch, Logstash, Kibana)
  • Proficiency in other monitoring tools (e.g., Prometheus, Grafana, Azure Monitor, App Insights, Splunk)
  • Experience with observability frameworks (metrics, distributed tracing, logging)
  • Experience working with cloud platforms (Azure preferred)
  • Strong scripting/programming skills (Python, Bash, etc.)
  • Understanding of distributed systems and microservices architecture
  • Solid understanding of security logging, audit trails, and system hardening
Additional Skills
  • Experience in tools such as Visual Studio, Azure DevOps, GitHub Enterprise, GitLab, CI/CD
  • Experience working in Financial Services / Insurance sector is an advantage
  • Experience building advanced automation scripts or tooling is a plus
  • Experience of working in an Agile environment and using Agile methodologies
What We Offer
  • Opportunity to work on enterprise‑scale, mission‑critical systems
  • Ownership of advanced observability and monitoring platforms
  • A culture of engineering excellence, automation, and continuous improvement
  • Collaboration with global teams and exposure to modern SRE practices
  • Continuous learning and professional growth opportunities
Benefits We Offer

Our benefits are designed to care for your holistic well‑being with programs for physical and mental health, financial wellness, and support for families.

We offer private health insurance for you and your family, life insurance, employer pension plan, meal and transportation allowance, as well as a work from home allowance. We also provide a cultural Heritage Day off, and “back to school” and “school report day” leaves and much more!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE: Observability & AI-Driven Reliability
SRE: Observability & AI-Driven Reliability

MetLife México • Fatih

On-site
TRY 320,000 - 540,000
Private health insurance
Pension plan
Work from home allowance
+1
Site Reliability Engineering Manager
Site Reliability Engineering Manager

n11 • Fatih

On-site
TRY 800,000 - 1,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems, Inc. • Turkey

On-site
TRY 600,000 - 900,000
Private health insurance
Continuous upskilling & development
English courses
+1
Site Reliability Engineer
Site Reliability Engineer

OBSS • Fatih

Hybrid
TRY 1,818,000 - 2,728,000
Flexible working arrangements
Training programs
Certifications
+1
Senior Site Reliability Engineer (Performance and Scalability)
Senior Site Reliability Engineer (Performance and Scalability)

JobCubby • Turkey

On-site
TRY 600,000 - 1,200,000
Immediate impact
Top compensation
Regional talent
SRE Leadership Manager - Scale & Reliability
SRE Leadership Manager - Scale & Reliability

n11 • Fatih

On-site
TRY 800,000 - 1,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Sezzle • Turkey

On-site
Innovative work environment
Competitive salary
Career advancement opportunities
Head of Platform & DevOps
Head of Platform & DevOps

RedCloud • Fatih

On-site
TRY 2,272,000 - 3,182,000
25 Days Annual leave
Enhanced Company Pension (Matched up to 5%)
Healthcare Cashplan
+3
Senior DevOps Engineer
Senior DevOps Engineer

WAGNIFY Bilgi Teknolojileri • Turkey

On-site
TRY 350,000 - 600,000
Senior SRE: GenAI-Driven Reliability & Cloud Ops
Senior SRE: GenAI-Driven Reliability & Cloud Ops

EPAM Systems, Inc. • Turkey

On-site
TRY 600,000 - 900,000
Private health insurance
Continuous upskilling & development
English courses
+1