Systems Reliability Engineering Senior Manager

Scotiabank

Bogotá

Presencial

COP 180.000.000 - 300.000.000

Jornada completa

Hace 13 días
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

ScotiaTech in Bogota is seeking an experienced Systems Reliability Engineering Senior Manager to lead the Level 5 Monitoring Team and ensure robust incident management, reporting, and continuous improvement of monitoring practices.

The role requires strong leadership, English/Spanish communication, and the ability to coordinate with cross-functional teams to maintain operational excellence in a fast-paced environment.

Formación

  • Bachelor's degree in Systems Engineering, Computer Science, Telecommunications, Technology Management, Industrial Engineering, or a related field.
  • Equivalent experience in critical IT operations may be considered.
  • Strong experience leading teams in operations, monitoring, NOC, SRE, incident management, or high-availability services.

Responsabilidades

  • Lead the daily operation of the Level 5 Monitoring Team, ensuring coverage and incident follow-up.
  • Manage shift schedules, rotations, and on-site/hybrid team continuity.
  • Oversee monitoring of applications and services across platforms and tools.
  • Escalate opportunities to improve observability, dashboards, and runbooks.
  • Prepare executive and operational reports on KPIs, capacity, and costs.

Conocimientos

Leadership
Executive communication
Stakeholder management
Prioritization
Team coordination

Educación

Bachelor's degree in Systems Engineering or related field
Equivalent experience in critical IT operations

Herramientas

Dynatrace
GEMS
Grail
ITSM

Descripción del empleo

Title: Systems Reliability Engineering Senior Manager

Thanks for your interest in ScotiaTech, Scotiabank's new and innovative Technology hub in Bogota.

Join a purpose driven winning team that promotes creativity and innovation in a fast-paced environment, where we’re always committed to results, in an inclusive, diverse, and high-performing culture.

Purpose

Contributes to the success of International Banking by leading the operation of the Monitoring HUB, ensuring that the Level 5 Analyst team executes monitoring, triage, alert handling, handovers, and escalations in accordance with established procedures. The role is responsible for operational coordination, administrative management, team performance, SLA/OLA tracking, financial control, executive reporting, and the continuous improvement of the monitoring model. All activities are carried out in compliance with applicable regulations, internal policies, control culture, operational risk management, security, compliance requirements, and standards of conduct.

Accountabilities

Lead the daily operation of the SRE Analyst Level 5 team, ensuring coverage, prioritization, operational discipline, incident follow-up, and adherence to established playbooks.

Manage shift scheduling, night rotation, extended coverage, backfills, vacations, absences, handovers, and the operational continuity of the monitoring team.

Define and maintain the team’s operational calendar, ensuring workload balance, assignment traceability, and alignment with business needs.

Oversee the monitoring of International Banking applications and services, ensuring that monitoring screens, alerts, dashboards, GEMS, Dynatrace, Grail, logs, and ITSM tools are reviewed according to the operating model.

Ensure alerts, service degradations, and incidents receive timely initial triage, sufficient evidence collection, preliminary classification, appropriate escalation, and follow-up through operational closure or handover.

Coordinate operational priorities with application, infrastructure, incident management, business operations, and SRE Central teams when support, decisions, or remediation fall outside the Level 5 scope.

Escalate opportunities and priorities to SRE Central related to observability, alerting, dashboards, thresholds, operational noise reduction, automation, remediation, playbooks, and regional standardization.

Propose improvements to alerting, dashboards, operational views, event correlation, service health monitoring, detection metrics, and procedures to reduce operational noise and improve MTTD, MTTR, and triage quality.

Govern the quality of shift handovers, operational logs, incident evidence, checklist compliance, and updates to operational documentation.

Manage team performance metrics, including shift coverage, handover compliance, alerts reviewed, incidents triaged, timely escalations, operational backlog, SLA/OLA adherence, documentation quality, and procedural compliance.

Prepare executive and operational reports on team performance, KPIs, capacity, allocation, costs, risks, alerting trends, recurring issues, and improvement initiatives.

Manage team costs, capacity planning, resource allocation, headcount requirements, productivity, operational budget, and forecasting required to sustain Monitoring HUB coverage.

Oversee administrative activities for the team, including onboarding, training, required access provisioning, knowledge development plans, compliance tracking, operational meetings, process feedback, and coordination of onsite activities.
Promote a strong control culture, compliance with procedures, operational risk management, information security, and responsible use of tools and access privileges.
Represent the monitoring team in operational forums, service reviews, incident committees, HUB governance meetings, and stakeholder follow-up sessions within International Banking.
Foster an inclusive, collaborative, and high-performing work environment focused on service excellence, continuous learning, operational reliability, and continuous improvement.

Education / Experience / Other Information
  • Bachelor's degree in Systems Engineering, Computer Science, Telecommunications, Technology Management, Industrial Engineering, or a related field. Equivalent experience in critical IT operations may be considered.
  • Strong experience leading teams in operations, monitoring, NOC, SRE, incident management, application support, production support, or high-availability technology services.
  • Experience managing shift schedules, coverage models, operational continuity, staffing, capacity planning, resource allocation, cost control, and on-site or hybrid teams.
  • Practical knowledge of observability and monitoring tools, including Dynatrace, dashboards, alerts, logs, metrics, traces, GEMS, Grail, ITSM, and comparable platforms.
  • Ability to interpret operational and executive metrics, identify trends, prioritize improvements, and present information clearly to both technical and non-technical stakeholders.
  • Knowledge of ITIL or equivalent processes, including incident management, problem management, escalation management, service reviews, operational governance, and continuous improvement.
  • Knowledge of Site Reliability Engineering (SRE) management practices, including SLIs, SLOs, error budgets, toil reduction, reliability reporting, observability standards, and operational readiness.
  • Experience proposing improvements to alerting, dashboards, playbooks, runbooks, recurring event management, thresholds, and triage procedures.
  • Strong leadership, executive communication, organizational, negotiation, prioritization, conflict resolution, decision-making, and follow-up skills.
  • English proficiency is preferred for communication with regional or global teams. Candidates without English proficiency must demonstrate the ability to perform the role in Spanish and leverage established channels for interactions requiring English.
  • Availability to work on-site in Bogotá and to support critical events or coordination needs outside standard business hours when required by the operating model.
Working Conditions

On-site role in Bogotá, Colombia, leading the Level 5 Monitoring Team from the office.
Standard schedule is Monday through Friday, from 8:00 a.m. to 5:30 p.m., with flexibility required to support operational reviews

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Systems Reliability Engineering Senior Manager
Systems Reliability Engineering Senior Manager

Scotiabank • Bogotá ciudad

Híbrido
COP 180.000.000 - 300.000.000
Senior SRE & Monitoring Hub Lead
Senior SRE & Monitoring Hub Lead

Scotiabank • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Senior Manager, Systems Reliability & Monitoring
Senior Manager, Systems Reliability & Monitoring

Scotiabank • Bogotá ciudad

Híbrido
COP 180.000.000 - 300.000.000
Senior Cloud Site Reliability Engineer Manager
Senior Cloud Site Reliability Engineer Manager

Scotiabank • Bogotá

Presencial
COP 300.000.000 - 420.000.000
Senior/Specialist SRE, Colombia
Senior/Specialist SRE, Colombia

CI&T • Colombia

Presencial
COP 72.000.000 - 96.000.000
Maternity and Parental leaves
Mobile services subsidy
Sick pay – Life insurance
+3
Director IB Client Services Engineering
Director IB Client Services Engineering

Scotiabank • Bogotá

Híbrido
COP 300.000.000 - 600.000.000
Service Reliability Engineer
Service Reliability Engineer

1083 Amadeus IT Group Colombia, S.A.S. • Colombia

Presencial
COP 182.089.000 - 254.926.000
Competitive remuneration
Vacation and holiday paid time off
Health insurances
+3
SAP Success Factors Platform and Support Lead
SAP Success Factors Platform and Support Lead

Scotiabank • Bogotá

Híbrido
COP 180.000.000 - 240.000.000
Supervisor, Monitoring
Supervisor, Monitoring

Scotiabank • Bogotá ciudad

Presencial
COP 120.000.000 - 180.000.000
SRE - Observability Engineer
SRE - Observability Engineer

T-mapp Jobs • Bogotá

Presencial
COP 180.000.000 - 240.000.000
Competitive salary
Comprehensive health benefits
Continuous learning & certifications
+1