Team Lead - Production Engineering

768 TP ICAP Management Services Ltd (Philippines Branch)

Taguig

On-site

PHP 1,200,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

768 TP ICAP Management Services Ltd (Philippines Branch) is seeking a Production Engineering Team Lead in Taguig. The role involves leading a team to ensure the performance and reliability of critical trading platforms, managing production issues, and collaborating with stakeholders.

The ideal candidate will have a degree or equivalent experience, strong AWS knowledge, and a solid background in incident management and market data flows.

Qualifications

  • Strong understanding of AWS operational environments.
  • Experience supporting business-critical front/mid-office applications.
  • Proven ability to work across multidisciplinary teams.

Responsibilities

  • Lead a team of Production Engineers for trading platforms.
  • Oversee and govern production issue management.
  • Act as the primary escalation point for major incidents.

Skills

AWS operational environments
Incident management
Market data flows
Customer communication

Education

Degree level education or equivalent experience

Tools

Grafana
CloudWatch
ELK
Splunk

Job description

Role Overview

The Production Engineering Team Lead is responsible for leading a team of Production Engineers and ensuring the stability, reliability, and operational performance of business‑critical trading platforms. This role combines hands‑on technical leadership with people management, prioritisation, and process ownership. The Team Lead acts as the primary escalation point, overseeing major incidents, change activity, and problem management while maintaining clear communication with engineering, platform, and business stakeholders.

Role Responsibilities
  • Oversee and govern production issue management across the team, acting as the senior escalation point for complex or high‑impact issues.
  • Own overall production health for the service area, ensuring monitoring coverage, alert quality, and operational standards are consistently met.
  • Ensure consistent operational readiness across platforms and regions through process ownership and team coordination.
  • Help set observability standards and priorities, ensuring the team delivers meaningful, actionable monitoring aligned with platform risk.
  • Ensure effective change governance, balancing delivery velocity with platform stability and risk management.
  • Coordinate major incidents at a leadership level, ensuring appropriate technical ownership, communication, and stakeholder management.
  • Own problem management outcomes, ensuring lessons learned are embedded into processes, tooling, and team priorities.
  • Define reliability improvement priorities and roadmap, aligning team effort with platform risk and business needs.
  • Act as the primary interface between production engineering, delivery teams, and senior stakeholders.
  • Prioritise and sponsor automation initiatives, ensuring team capacity is focused on the highest operational value.
Experience and Competences
  • Essential – Educated to degree level or equivalent combination of education and experience.
  • Solid understanding of AWS operational environments, including load balancers, regional fail‑over behaviour, instance lifecycles, and managed databases.
  • Experience supporting business‑critical front/mid‑office applications.
  • Deep knowledge of market data flows, instrument definitions, pricing mechanisms, and session‑based connectivity.
  • Ability to interpret complex application logs and diagnose backend issues with accuracy and speed.
  • Strong root‑cause analysis capability and ability to evaluate symptoms, isolate faults, and determine remediation paths.
  • Solid experience in incident management, major‑incident coordination, and structured problem‑solving.
  • Demonstrated ability to work across regions, managing concurrent issues, escalations, and stakeholder communications.
  • Clear understanding of change‑management disciplines, including risk assessment and deployment validation.
  • Familiarity with observability tooling (Grafana, CloudWatch, ELK, Splunk), including metrics, logs, dashboards, and alerting used to assess system health and reliability.
  • Proven ability to work across multidisciplinary teams (Business, Operations, Developers, DevOps).
  • Strong customer‑focus and ability to communicate complex technical issues in a business‑friendly manner.
  • Comfortable supporting global operations and adapting to multi‑region workflows.
  • Demonstrated interest or experience in applying SRE principles such as reliability metrics, automation, and continuous improvement within a support or operations role.
  • Experience contributing to improved mean time to detect (MTTD) and mean time to restore (MTTR) through better observability, tooling, or process.
  • Understanding of the balance between feature delivery and operational stability in business‑critical systems.
Desired
  • Solid experience supporting trading platforms, financial exchanges, or real‑time transactional systems.
  • Reasonable exposure to FIX‑based workflows, messaging pipelines, or market‑connectivity architectures.
  • Experience working within AWS‑native or hybrid‑cloud financial environments.
  • Reasonable experience collaborating with front‑office trading desks or broker support teams.
  • Familiarity with CI/CD pipelines, DevOps practices, or automated deployment frameworks.
  • Exposure to SRE concepts such as SLIs, SLOs, error budgets, or reliability metrics within a support or operations context.
  • Experience improving platform resilience through automation, monitoring enhancements, or operational tooling.
  • Familiarity with failure scenarios, recovery patterns, or high‑availability strategies in distributed systems.
Location

Philippines – Ecoprime Building, Taguig City

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Production Engineer
Senior Production Engineer

863 Parameta Solutions (Singapore) Pte. Limited • Taguig

On-site
PHP 558,000 - 781,200
Production Engineer
Production Engineer

768 TP ICAP Management Services Ltd (Philippines Branch) • Taguig

On-site
PHP 446,400 - 669,600
Team Lead - Production Engineering
Team Lead - Production Engineering

TP ICAP • Manila

On-site
PHP 1,800,000 - 2,400,000
Sr. Production Engineer (Trading Applications) – Hybrid | Taguig City, Philippines
Sr. Production Engineer (Trading Applications) – Hybrid | Taguig City, Philippines

Charterhouse Pte Ltd • Taguig

On-site
PHP 600,000 - 800,000
Production Engineer
Production Engineer

ECLARO • Taguig

On-site
PHP 900,000 - 1,200,000
Production Engineering Team Lead: Reliability & Incident Response
Production Engineering Team Lead: Reliability & Incident Response

768 TP ICAP Management Services Ltd (Philippines Branch) • Taguig

On-site
PHP 1,200,000 - 1,500,000
Production Engineering Lead — Reliability & Incidents
Production Engineering Lead — Reliability & Incidents

TP ICAP Group • Taguig

Hybrid
PHP 1,200,000 - 1,500,000
Team Lead - Production Engineering
Team Lead - Production Engineering

TP ICAP Group • Taguig

On-site
PHP 1,200,000 - 1,500,000
Production Engineer, SRE‑Driven Trading Platform
Production Engineer, SRE‑Driven Trading Platform

768 TP ICAP Management Services Ltd (Philippines Branch) • Taguig

On-site
PHP 446,400 - 669,600
Senior Production Engineer
Senior Production Engineer

TP ICAP Group • Taguig

On-site
PHP 1,200,000 - 1,500,000
Inclusive work environment
Opportunities for growth
Employee networking programs