Technical Service Operations Lead (TSO Lead), Kuala Lumpur

Xsolla

Kuala Lumpur

On-site

MYR 240,000 - 300,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance (employee and dependants)
Flexible hours
Free trainings and specialized conferences participation
Convenient work tools

Job summary

A leading gaming technology company in Kuala Lumpur seeks a Technical Service Operations Lead (TSO Lead) to oversee incident management within their Global Technical Operations team. This role involves serving as an Incident Commander for major incidents, managing communications, and analyzing incident trends. Candidates should have 6+ years of experience in technical operations, strong ITIL knowledge, and excellent communication skills. The position offers a salary between RM240,000 and RM300,000 a year.

Qualifications

  • 6+ years of experience in incident management, SRE, NOC leadership, or technical operations.
  • Proven incident management experience coordinating multi-team responses.
  • Strong ITIL foundation with practical experience implementing workflows.

Responsibilities

  • Serve as Incident Commander for major incidents, coordinating cross-functional teams.
  • Own incident communications, updating stakeholders throughout the incident lifecycle.
  • Facilitate blameless Post-Incident Reviews for major incidents.

Skills

Incident management
Technical operations
ITIL knowledge
Observability/monitoring expertise
Analytical mindset
Excellent communication

Education

6+ years in incident management or technical operations

Tools

Datadog
JIRA Service Management
PagerDuty
Slack
Confluence
OpsGenie

Job description

About You

We are looking for a Technical Service Operations Lead (TSO Lead) who is operationally driven, collaborative, analytical, and a strong communicator to join our Global Technical Operations (GTO) team. The best candidate will thrive in a fast‑paced, highly collaborative and exceptionally dynamic setting and be excited to help coordinate incident response alongside cross‑functional teams, identify trends and patterns in production issues, improve how we communicate with partners during incidents, and drive continuous improvement through post‑incident reviews.

Strong incident management experience, ITIL knowledge, and observability/monitoring expertise are essential, along with experience in technical operations, SRE, or NOC environments supporting high‑availability platforms (payments, e‑commerce, SaaS, or gaming). The ability to communicate clearly and effectively in English — both written and verbal — across technical and executive audiences will be key to your success in this role.

If you’re passionate about driving operational excellence and platform reliability at scale and love ensuring the reliability and uptime of commerce and payment solutions that game developers and players depend on, we would love to hear from you!

Responsibilities
  • Serve as Incident Commander for major incidents — coordinating cross‑functional response teams, driving investigation, making escalation decisions, and ensuring incidents are resolved within SLA targets.
  • Own all incident communications: draft and send clear, timely updates to senior leadership, Customer Success, and partner/customer contacts throughout the incident lifecycle, and manage customer‑facing status page updates (status.xsolla.com).
  • Facilitate blameless Post‑Incident Reviews (PIRs) for major incidents — leading root cause identification, assigning corrective actions with clear owners and deadlines, and tracking them to closure.
  • During non‑incident periods, proactively analyze incident trends, recurring issues, and production bugs — identify patterns, create Problem tickets, and report findings and recommendations to product and engineering teams on a regular cadence.
  • Enforce the incident management framework across the organization, including the severity model, priority matrix, SLA targets, escalation procedures, and deployment readiness gates.
  • Oversee and mentor the Operations Engineer on your shift — coaching on triage, investigation, runbook execution, and documentation quality while conducting regular knowledge transfer sessions to build depth across the service portfolio.
  • Produce shift handoff reports and deliver regular operational reporting: incident trends, KPI performance (MTTD, MTTA, MTTR), SLA adherence, proactive detection rate, and repeat incident analysis.
  • Audit service catalogue completeness on a regular cadence and govern JIRA Service Management workflows for incident, PIR, and problem management.
  • Cover for the Operations Engineer role during absences, breaks, or surge incidents. Participate in weekend on‑call rotation for major incidents.
Qualifications
  • 6+ years of experience in incident management, SRE, NOC leadership, or technical operations in a production environment supporting high‑availability, high‑transaction systems (payments, e‑commerce, SaaS, or gaming platforms preferred).
  • Proven incident management experience — coordinating multi‑team response, making real‑time escalation decisions, and communicating with executive stakeholders under pressure.
  • Excellent written and verbal communication skills in English — ability to draft clear, concise executive updates at 3 AM under pressure, facilitate blameless PIRs, present operational metrics to senior leadership, and communicate incident status to customers and partners with clarity and professionalism.
  • Strong ITIL foundation — understanding of incident, problem, and change management lifecycles with practical experience implementing or operating ITIL‑aligned workflows.
  • Technical depth across the observability stack — ability to read and interpret logs, traces, and metrics in Datadog (or equivalent: Grafana, Splunk, New Relic). Understanding of APM, SLOs, error budgets, burn‑rate alerting, and synthetic monitoring.
  • Hands‑on experience with incident tooling: Datadog, PagerDuty or OpsGenie, JIRA or JIRA Service Management, Slack, and Confluence.
  • Analytical mindset — ability to identify trends, patterns, and recurring issues from incident data and translate them into actionable recommendations for product and engineering teams.
  • Experience with SLA/SLO‑driven operations where MTTD, MTTA, and MTTR are measured, reported, and improved.
  • Experience with or strong interest in AI/ML‑assisted operations: anomaly detection, alert correlation, predictive alerting, automated remediation, or self‑healing automation.
  • Comfort with 24×7 shift‑based operations as part of a follow‑the‑sun model with handoff overlaps. Weekend on‑call (rotating) for critical severities is required.
Nice to have
  • Experience in the gaming, payments, or fintech industry.
  • Experience with customer/partner‑facing incident communications and status page management.
  • JIRA Service Management administration experience: workflows, SLA timers, automation rules, queues, and permissions.
  • Familiarity with Datadog Service Catalog, scorecards, and SLOs — especially burn‑rate alerts and multi‑window SLOs.
  • Experience building an operations function from scratch — defining processes, writing runbooks, establishing governance cadences.
  • Background in Kubernetes, cloud infrastructure (GCP preferred), microservices architecture, or distributed systems.
  • ITIL certification (Foundation or higher) is a plus but not required.
Salary

RM240,000 - RM300,000 a year

Benefits
  • Convenient work tools: Latest Mac workplaces plus additional hardware to make you more effective at work.
  • Google Chat, Gmail, Google Drive, Confluence, Jira, GitLab.
  • Professional growth: Free trainings and participation in specialized conferences; rich knowledge exchange within the company.
  • Health insurance (Medical, dental and optical) – Employee and dependants; flexible hours; no dress code; comfortable and new office environment.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Service Operations Lead (TSO Lead), Kuala Lumpur
Technical Service Operations Lead (TSO Lead), Kuala Lumpur

Xsolla • Kuala Lumpur

On-site
MYR 100,000 - 130,000
Health insurance
Flexible hours
Free trainings and conferences
+1
Global Technical Operations Lead - Incident & Reliability
Global Technical Operations Lead - Incident & Reliability

Xsolla • Kuala Lumpur

On-site
MYR 100,000 - 130,000
Health insurance
Flexible hours
Free trainings and conferences
+1
Managed Services
Managed Services

Doherty • Kuala Lumpur

On-site
MYR 80,000 - 110,000
Competitive salary plus performance-related bonus
Hybrid working
Day off on your birthday
+4
Senior Cloud And Infrastructure Engineer (Level 3)
Senior Cloud And Infrastructure Engineer (Level 3)

Doherty • Kuala Lumpur

On-site
MYR 180,000 - 260,000
Competitive salary
Performance-related bonus
Hybrid working
+4
Senior Specialist, SRE
Senior Specialist, SRE

TNG Digital • Kuala Lumpur

On-site
MYR 240,000 - 360,000
Medical coverage
Extra leave
Lifestyle allowance
+2
Senior Operations Manager, Global Service Support Centre
Senior Operations Manager, Global Service Support Centre

TNS Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Security Operations Center Lead
Security Operations Center Lead

Altera • Bayan Lepas

On-site
MYR 180,000 - 280,000
Senior Associate, Incident Response
Senior Associate, Incident Response

S-RM • Malaysia

Hybrid
MYR 180,000 - 280,000
20 days paid holiday
Flexible working
EPF pension
+3
NOC Engineer
NOC Engineer

Soprano Design • Kuala Lumpur

On-site
MYR 33,000 - 56,000
Hybrid working
Birthday leave
Employee assistance program
+2
Specialist, Incident Management
Specialist, Incident Management

TNG Digital • Kuala Lumpur

On-site
MYR 180,000 - 260,000
Medical coverage
Extra leave for family and caregiving
Monthly lifestyle allowance via TNG eW
+2