Technical Lead

HCL Technologies Limited

Almoloya de Juárez

Presencial

MXN 900.000 - 1.300.000

Jornada completa

Hace 4 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

HCL Technologies Limited is seeking a Site Reliability Engineer to drive deployment readiness and maintain operational stability for payment platforms. You will implement automation, optimize monitoring, and lead incident management within a fast-paced, cross-functional team.

The role emphasizes production readiness and continuous improvement across DevOps practices. The position interfaces with engineering, infrastructure and external partners to ensure reliable service delivery, effective

Formación

  • Experience in Site Reliability Engineering or similar roles.
  • Strong incident management and root-cause analysis abilities.
  • Hands-on troubleshooting across apps, infra, databases and networks.

Responsabilidades

  • Manage end-to-end platform availability and performance.
  • Lead incident triage, remediation, and communications with stakeholders.
  • Support change and release processes for production deployments.
  • Improve monitoring, alerts, and dashboards to reduce toil.
  • Collaborate with cross-functional teams and drive automation initiatives.

Conocimientos

Incident management
Monitoring
Automation
DevOps
Communication

Herramientas

Observability tools
CI/CD pipelines

Descripción del empleo

Role OverviewThe Real Time Payments International team is looking for a Site Reliability Engineer (SRE) to drive application deployment readiness, manage day-to-day operational stability and support the reliability of critical payment platforms by implementing automation, leverage best practices and work with a high‑impact team responsible for driving production readiness, reliability, and DevOps automation across Mastercard platforms.This role plays a key part in incident management, change readiness, and platform operations, while contributing to continuous improvement initiatives.Key ResponsibilitiesPlatform Operations & StabilitySupport end-to-end availability, monitoring, and performance of critical payment platforms.Execute operational processes to ensure platform health and stability.Participate in capacity checks, readiness validations, and environment monitoring.Incident Management & ExecutionActively manage and coordinate incident triage and resolution.Serve as incident commander driving medium to high-severity incidents.Ensure timely updates, accurate impact assessment, and appropriate escalation.Contribute to root cause analysis with clear identification of actions and ownership.Change & Release SupportParticipate in highlighting gaps and defining test cases required for a change in lower environments and validate lower environment test completeness.Ensure adherence to change governance processes (test case reviews, checklists, approvals, rollback readiness).Engage in creating change plans and support execution of production changes, deployments, and validations.Technical TroubleshootingPerform hands-on troubleshooting across:Application behaviour and dependencies.Infrastructure components (compute, network, storage).Database and performance issues.Collaborate with engineering, infrastructure and other technical teams to isolate and resolve issues efficiently.Monitoring & ObservabilityImprove system health monitoring using observability tools and alerts.Identify gaps in alerting and contribute to improving quality of alerting and dashboards.Ensure proactive detection of anomalies using observability tools.Automation & Process ImprovementContribute to automation initiatives to reduce toil and errors.Identify repetitive operational tasks and drive improvements.Support implementation of DevOps best practices.Leverage AI-driven tools to improve monitoring, incident detection, and operational efficiency, enabling faster troubleshooting and reduced manual effort in day-to-day operations.Stakeholder CoordinationWork closely with engineering, program teams, and external partners during incidents and changes.Provide structured updates to stakeholders with clarity and consistency.Ensure alignment during critical activities.Risk IdentificationHighlight operational and platform risks including test coverage gaps, infrastructure constraints, dependency risks.Escalate issues proactively and support mitigation tracking.Team Contribution & MentorshipSupport onboarding and guidance of junior team members.Contribute to runbooks, documentation, and knowledge sharing.Drive consistency in execution and adherence to operational standards.

Key Responsibilities

Role OverviewThe Real Time Payments International team is looking for a Site Reliability Engineer (SRE) to drive application deployment readiness, manage day-to-day operational stability and support the reliability of critical payment platforms by implementing automation, leverage best practices and work with a high‑impact team responsible for driving production readiness, reliability, and DevOps automation across Mastercard platforms.This role plays a key part in incident management, change readiness, and platform operations, while contributing to continuous improvement initiatives.Key ResponsibilitiesPlatform Operations & StabilitySupport end-to-end availability, monitoring, and performance of critical payment platforms.Execute operational processes to ensure platform health and stability.Participate in capacity checks, readiness validations, and environment monitoring.Incident Management & ExecutionActively manage and coordinate incident triage and resolution.Serve as incident commander driving medium to high-severity incidents.Ensure timely updates, accurate impact assessment, and appropriate escalation.Contribute to root cause analysis with clear identification of actions and ownership.Change & Release SupportParticipate in highlighting gaps and defining test cases required for a change in lower environments and validate lower environment test completeness.Ensure adherence to change governance processes (test case reviews, checklists, approvals, rollback readiness).Engage in creating change plans and support execution of production changes, deployments, and validations.Technical TroubleshootingPerform hands‑on troubleshooting across:Application behaviour and dependencies.Infrastructure components (compute, network, storage).Database and performance issues.Collaborate with engineering, infrastructure and other technical teams to isolate and resolve issues efficiently.Monitoring & ObservabilityImprove system health monitoring using observability tools and alerts.Identify gaps in alerting and contribute to improving quality of alerting and dashboards.Ensure proactive detection of anomalies using observability tools.Automation & Process ImprovementContribute to automation initiatives to reduce toil and errors.Identify repetitive operational tasks and drive improvements.Support implementation of DevOps best practices.Leverage AI-driven tools to improve monitoring, incident detection, and operational efficiency, enabling faster troubleshooting and reduced manual effort in day‑to‑day operations.Stakeholder CoordinationWork closely with engineering, program teams, and external partners during incidents and changes.Provide structured updates to stakeholders with clarity and consistency.Ensure alignment during critical activities.Risk IdentificationHighlight operational and platform risks including test coverage gaps, infrastructure constraints, dependency risks.Escalate issues proactively and support mitigation tracking.Team Contribution & MentorshipSupport onboarding and guidance of junior team members.Contribute to runbooks, documentation, and knowledge sharing.Drive consistency in execution and adherence to operational standards.

Skill Requirements

Role OverviewThe Real Time Payments International team is looking for a Site Reliability Engineer (SRE) to drive application deployment readiness, manage day‑to‑day operational stability and support the reliability of critical payment platforms by implementing automation, leverage best practices and work with a high‑impact team responsible for driving production readiness, reliability, and DevOps automation across Mastercard platforms.This role plays a key part in incident management, change readiness, and platform operations, while contributing to continuous improvement initiatives.Key ResponsibilitiesPlatform Operations & StabilitySupport end‑to‑end availability, monitoring, and performance of critical payment platforms.Execute operational processes to ensure platform health and stability.Participate in capacity checks, readiness validations, and environment monitoring.Incident Management & ExecutionActively manage and coordinate incident triage and resolution.Serve as incident commander driving medium to high‑severity incidents.Ensure timely updates, accurate impact assessment, and appropriate escalation.Contribute to root cause analysis with clear identification of actions and ownership.Change & Release SupportParticipate in highlighting gaps and defining test cases required for a change in lower environments and validate lower environment test completeness.Ensure adherence to change governance processes (test case reviews, checklists, approvals, rollback readiness).Engage in creating change plans and support execution of production changes, deployments, and validations.Technical TroubleshootingPerform hands‑on troubleshooting across:Application behaviour and dependencies.Infrastructure components (compute, network, storage).Database and performance issues.Collaborate with engineering, infrastructure and other technical teams to isolate and resolve issues efficiently.Monitoring & ObservabilityImprove system health monitoring using observability tools and alerts.Identify gaps in alerting and contribute to improving quality of alerting and dashboards.Ensure proactive detection of anomalies using observability tools.Automation & Process ImprovementContribute to automation initiatives to reduce toil and errors.Identify repetitive operational tasks and drive improvements.Support implementation of DevOps best practices.Leverage AI‑driven tools to improve monitoring, incident detection, and operational efficiency, enabling faster troubleshooting and reduced manual effort in day‑to‑day operations.Stakeholder CoordinationWork closely with engineering, program teams, and external partners during incidents and changes.Provide structured updates to stakeholders with clarity and consistency.Ensure alignment during critical activities.Risk IdentificationHighlight operational and platform risks including test coverage gaps, infrastructure constraints, dependency risks.Escalate issues proactively and support mitigation tracking.Team Contribution & MentorshipSupport onboarding and guidance of junior team members.Contribute to runbooks, documentation, and knowledge sharing.Drive consistency in execution and adherence to operational standards.

Other Requirements

Role OverviewThe Real Time Payments International team is looking for a Site Reliability Engineer (SRE) to drive application deployment readiness, manage day‑to‑day operational stability and support the reliability of critical payment platforms by implementing automation, leverage best practices and work with a high‑impact team responsible for driving production readiness, reliability, and DevOps automation across Mastercard platforms.This role plays a key part in incident management, change readiness, and platform operations, while contributing to continuous improvement initiatives.Key ResponsibilitiesPlatform Operations & StabilitySupport end‑to‑end availability, monitoring, and performance of critical payment platforms.Execute operational processes to ensure platform health and stability.Participate in capacity checks, readiness validations, and environment monitoring.Incident Management & ExecutionActively manage and coordinate incident triage and resolution.Serve as incident commander driving medium to high‑severity incidents.Ensure timely updates, accurate impact assessment, and appropriate escalation.Contribute to root cause analysis with clear identification of actions and ownership.Change & Release SupportParticipate in highlighting gaps and defining test cases required for a change in lower environments and validate lower environment test completeness.Ensure adherence to change governance processes (test case reviews, checklists, approvals, rollback readiness).Engage in creating change plans and support execution of production changes, deployments, and validations.Technical TroubleshootingPerform hands‑on troubleshooting across:Application behaviour and dependencies.Infrastructure components (compute, network, storage).Database and performance issues.Collaborate with engineering, infrastructure and other technical teams to isolate and resolve issues efficiently.Monitoring & ObservabilityImprove system health monitoring using observability tools and alerts.Identify gaps in alerting and contribute to improving quality of alerting and dashboards.Ensure proactive detection of anomalies using observability tools.Automation & Process ImprovementContribute to automation initiatives to reduce toil and errors.Identify repetitive operational tasks and drive improvements.Support implementation of DevOps best practices.Leverage AI‑driven tools to improve monitoring, incident detection, and operational efficiency, enabling faster troubleshooting and reduced manual effort in day‑to‑day operations.Stakeholder CoordinationWork closely with engineering, program teams, and external partners during incidents and changes.Provide structured updates to stakeholders with clarity and consistency.Ensure alignment during critical activities.Risk IdentificationHighlight operational and platform risks including test coverage gaps, infrastructure constraints, dependency risks.Escalate issues proactively and support mitigation tracking.Team Contribution & MentorshipSupport onboarding and guidance of junior team members.Contribute to runbooks, documentation, and knowledge sharing.Drive consistency in execution and adherence to operational standards.

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry‑leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026totaled $14.8billion.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior SRE Lead – Real-Time Payments Platform
Senior SRE Lead – Real-Time Payments Platform

HCL Technologies Limited • Almoloya de Juárez

Presencial
MXN 900.000 - 1.300.000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Oracle • Región Centro

Presencial
MXN 600.000 - 900.000
Senior Technical Lead
Senior Technical Lead

HCL Technologies Limited • Aguascalientes

Presencial
MXN 500.000 - 700.000
SRE (Engineering & Administration Background)
SRE (Engineering & Administration Background)

fulcrumdigital • Ciudad de México

Híbrido
MXN 900.000 - 1.500.000
Software Development Engineer III IRC301892
Software Development Engineer III IRC301892

GlobalLogic • México

Presencial
MXN 900.000 - 1.200.000
Exciting Projects
Collaborative Environment
Work-Life Balance
+2
Technology Lead Backend - Developer (Java, Microservices, Kafka, Couchbase)
Technology Lead Backend - Developer (Java, Microservices, Kafka, Couchbase)

Infosys Limited • Región Centro

Presencial
MXN 600.000 - 1.200.000
Customer Success Associate IRC301486
Customer Success Associate IRC301486

GlobalLogic • Región Centro

Presencial
MXN 180.000 - 300.000
Flexible work schedules
English classes
International travel opportunities
+1
Senior Site Reliability Engineer - Scale, Automate, Resilient Ops
Senior Site Reliability Engineer - Scale, Automate, Resilient Ops

Mastercard • Ciudad de México

Presencial
MXN 1.200.000 - 2.000.000
Agentic Solutions Specialist
Agentic Solutions Specialist

Mastercard • Estado de México

Presencial
MXN 900.000 - 1.300.000
HPE Networking SE MTY
HPE Networking SE MTY

Hewlett Packard Enterprise Development LP • Ciudad de México

Híbrido
MXN 1.200.000 - 1.800.000