Senior Site Reliability & Observability Engineer (SRE)

TransUnion

Mexico

On-site

PHP 600,000 - 900,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TransUnion’s Reliability Engineering team ensures the stability, availability, and performance of mission-critical platforms. The role focuses on observability, incident management, and platform engineering to support business operations.

This on-site position requires in-person work at a TU office location. The candidate will lead incident responses, perform RCAs, and define SLIs/SLOs to drive continuous improvement.

Qualifications

  • Bachelor’s degree in Computer Science, Systems Engineering, Telecommunications Engineering, or a related technical discipline.
  • Proven experience in Site Reliability Engineering (SRE) or Production Operations.
  • Strong knowledge of distributed systems, observability practices, incident management, and troubleshooting methodologies.
  • Experience managing major incidents and root cause analysis processes.
  • Understanding of reliability frameworks including SLOs, SLIs, and error budgets.

Responsibilities

  • Ensure the reliability, stability, and availability of mission-critical services through SRE best practices.
  • Operate and improve the observability platform, including metrics, logs, traces, and alerting.
  • Monitor critical systems and respond to incidents to minimize service disruption.
  • Lead major incident response activities and coordinate recovery efforts and communications.
  • Conduct RCA and blameless post-mortems to identify improvements.
  • Define and monitor SLIs, SLOs, and error budgets.
  • Automate processes to reduce toil through scripting and platform engineering.
  • Support and troubleshoot distributed environments including Cassandra, Kafka, Kubernetes, and databases.
  • Collaborate with teams to improve deployment reliability and platform performance.
  • Drive continuous improvement in monitoring, availability, and resilience.

Skills

SRE
Observability
Incident management

Education

Bachelor's degree in Computer Science

Tools

Apache Cassandra
Kafka
Kubernetes

Job description

TransUnion's Job Applicant Privacy Notice Team OverviewThe Reliability Engineering team ensures the stability, availability, and performance of Buró de Crédito’s mission-critical platforms and services. Through observability, automation, and incident management practices, the team drives operational excellence and continuous improvement. Working closely with Infrastructure, Development, Database, and Security teams, they help maintain resilient systems that support critical business operations. This job is assigned as On-Site Essential and requires in- person work at an assigned TU office location as a condition of employment.

Role Overview And Core Responsibilities
  • Ensure the reliability, stability, and availability of mission-critical services through Site Reliability Engineering (SRE) best practices.
  • Operate and continuously improve the organization’s observability platform, including metrics, logs, traces, and alerting capabilities.
  • Monitor critical systems proactively and respond to operational incidents to minimize service disruption and business impact.
  • Lead major incident response activities, coordinating recovery efforts and stakeholder communication during service outages.
  • Conduct root cause analysis (RCA) and facilitate blameless post-mortems to identify systemic improvements and prevent recurrence.
  • Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Automate operational processes and reduce manual effort (toil) through scripting, infrastructure automation, and platform engineering practices.
  • Support and troubleshoot distributed environments including Apache Cassandra, Kafka, Kubernetes, relational databases, and NoSQL platforms.
  • Collaborate with development and infrastructure teams to improve deployment reliability, operational readiness, and platform performance.
  • Drive continuous improvement initiatives focused on monitoring, availability, scalability, resilience, and operational efficiency.
Required Knowledge And Experiences
  • Bachelor’s degree in Computer Science, Systems Engineering, Telecommunications Engineering, or a related technical discipline, providing the foundation required to manage complex distributed environments.
  • Proven experience in Site Reliability Engineering (SRE), Production Operations, Platform Engineering, Infrastructure Engineering, or Reliability-focused roles.
  • Strong knowledge of distributed systems, observability practices, incident management, and troubleshooting methodologies for mission-critical environments.
  • Experience managing major incidents, root cause analysis processes, service restoration activities, and operational excellence initiatives.
  • Understanding of reliability frameworks including SLOs, SLIs, error budgets, continuous improvement, and service management best practices.
#LI-SG4 TransUnion Overview:

At TransUnion, we encourage and are committed to creating a real, positive impact and shared sense of purpose within our Workforce for Good, which empowers our people to grow, innovate and contribute to a better future for our communities and customers. We strive to build an environment where our associates are in the driver’s seat of their professional development— while having access to help along the way. We recognize that success comes when our associates thrive both professionally and personally; that’s why we prioritize work/life flexibility and offer resources for our teams across the globe to collaborate and drive excellence. Be a part of our Workforce for Good – you’ll work with great people, pioneering products and cutting-edge technology.

TransUnion is a global information and insights company with over 12,000 associates operating in more than 30 countries. We make trust possible by ensuring each person is reliably represented in the marketplace. We do this with a Tru™ picture of each person: an actionable view of consumers, stewarded with care. Through our acquisitions and technology investments we have developed innovative solutions that extend beyond our strong foundation in core credit into areas such as marketing, fraud, risk and advanced analytics. As a result, consumers and businesses can transact with confidence and achieve great things. We call this Information for Good® — and it leads to economic opportunity, great experiences and personal empowerment for millions of people around the world.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE & Observability Engineer: Incident & Reliability
Senior SRE & Observability Engineer: Incident & Reliability

TransUnion • Mexico

Hybrid
PHP 600,000 - 900,000
Content & Social Media
Content & Social Media

TransUnion • Mexico

Hybrid
PHP 1,275,000 - 1,912,000
Hybrid work model
Advisor, Fraud Solutions Consulting MX
Advisor, Fraud Solutions Consulting MX

TransUnion • Mexico

Hybrid
PHP 3,183,000 - 4,599,000
Core Products - Product Manager MX
Core Products - Product Manager MX

TransUnion • Mexico

Hybrid
PHP 2,476,000 - 3,538,000
Fraud Solutions Delivery Lead - LATAM (Hybrid)
Fraud Solutions Delivery Lead - LATAM (Hybrid)

TransUnion • Mexico

Hybrid
PHP 4,923,000 - 7,385,000
Fraud Product Delivery (Product Owner) MX
Fraud Product Delivery (Product Owner) MX

TransUnion • Mexico

Hybrid
PHP 4,923,000 - 7,385,000
Hybrid work model
Major Account Executive Banking
Major Account Executive Banking

TransUnion • Mexico

On-site
PHP 5,180,000 - 7,313,000
Customer Service Analyst
Customer Service Analyst

TransUnion • Mexico

Hybrid
PHP 2,344,000 - 3,208,000
Sales Specialist, Fraud Solutions
Sales Specialist, Fraud Solutions

TransUnion • Makati

Hybrid
PHP 1,800,000 - 3,200,000
Senior SRE & Observability Engineer — Reliability & Automation
Senior SRE & Observability Engineer — Reliability & Automation

TransUnion • Mexico

Hybrid
PHP 3,663,000 - 5,495,000