Incident Management - Lead TechOps Engineer

Bybit

Kuala Lumpur

On-site

MYR 120,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Study Growth Fund
Internal Events
Global Collaboration
Career Advancement
Internal Mobility

Job summary

Bybit, a leading cryptocurrency exchange, is seeking an experienced Incident Response Lead in Kuala Lumpur to join our global Ops team. You will own 24/7 on-call coordination, drive recovery, and ensure rapid resolution for major production incidents across our fintech platforms.

You will align cross-functional teams, produce post-mortems, and build robust monitoring and runbooks to prevent recurrence. English fluency and a proven ability to operate with influence rather than authority are

Qualifications

  • Experience independently leading major incident response, including on-site coordination, management reporting, post-mortems, and risk closure.
  • Ability to coordinate multiple technical and business teams without direct authority.
  • Strong technical comprehension — able to understand system architecture, service dependencies, monitoring/alerting, and incident chains.
  • Strong written communication skills — able to produce well-structured post-mortems with factual evidence and clear conclusions.
  • Experience with large-scale internet, fintech, payments, trading platforms, or other high-availability systems.
  • Working proficiency in English — able to participate in English meetings and handle routine written communication.
  • Willingness to accept 24/7 on-call rotation.

Responsibilities

  • Participate in 24/7 on-call rotation and respond to major production incidents.
  • Serve as incident coordination hub, assemble teams, chair emergency meetings, and drive business recovery.
  • Determine incident severity and escalation timing, keeping management informed of incident impact and next steps.
  • Unify the factual narrative and coordinate with PR and CS on external communications.
  • Lead incident post-mortems: reconstruct timelines, document impact, actions, and decisions.
  • Organize root cause analysis and drive improvements in monitoring and processes.
  • Track action items to on-time completion and verify effectiveness through testing or data.
  • Regularly produce reports on incident trends, recurring issues, and major risks.

Skills

Incident management
Cross-functional coordination
Technical understanding
Communications
Leadership without authority

Tools

PagerDuty
Datadog
Grafana

Job description

Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance.

Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution.

Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services.

Job Responsibilities
  • Participate in 24/7 on-call rotation, responding promptly to major production incidents
  • Serve as the incident coordination hub, assembling relevant teams (R&D, SRE, Security, Risk, PR, CS, etc.), chairing emergency meetings, and driving business recovery
  • Determine incident severity and escalation timing, keeping management informed of incident impact, response progress, recovery expectations, and next steps
  • Unify the factual narrative when information is incomplete, and coordinate with PR and CS on external communications
  • Lead incident post-mortems: reconstruct timelines, document incident impact, response actions, and key decisions
  • Organize technical teams to complete root cause analysis, identify gaps in monitoring/alerting, system resilience, change management, and processes; drive incident classification, accountability assignment, and action item confirmation
  • Track action items to on-time completion, verifying effectiveness through testing, drills, or monitoring data
  • Regularly produce reports on incident trends, recurring issues, and major risks
Process & Capability Building
  • Establish and continuously optimize incident response SOPs, Runbooks, escalation paths, and compliance reporting processes; maintain on-call schedules and escalation chains
  • Organize emergency response training and cross-department drills; drive incident management tooling, data dashboards, and automation capabilities
Who We're Looking For
  • Experience independently leading major incident response, including on-site coordination, management reporting, post-mortems, and risk closure
  • Ability to coordinate multiple technical and business teams without direct authority, continuously driving resolution under pressure
  • Strong technical comprehension — able to understand system architecture, service dependencies, monitoring/alerting, and incident chains, and judge whether root cause analysis and remediation plans are complete
  • Ability to quickly distill key information, deliver concise and clear briefings to management, take ownership of outcomes, and follow through until issues are closed
  • Strong written communication skills — able to produce well-structured post-mortems with factual evidence and clear conclusions
  • Experience with large-scale internet, fintech, payments, trading platforms, or other high-availability systems
  • Working proficiency in English — able to participate in English meetings and handle routine written communication
  • Willingness to accept 24/7 on-call rotation
Nice-to-Haves
  • Experience in crypto, exchanges, payments, or financial trading systems
  • Experience in SRE, production operations, technical support, or reliability engineering
  • Familiarity with PagerDuty, Datadog, Grafana, or similar monitoring and alerting tools
  • Familiarity with ITIL, Incident Management, or Problem Management frameworks
Why Join Us

At Bybit, we are committed to fostering a supportive and enriching work environment.

Our benefits include:

  • Study Growth Fund: We support your professional development and continuous learning.
  • Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation.
  • Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world.
  • Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company.
  • Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead TechOps Engineer New Kuala Lumpur, Malaysia
Lead TechOps Engineer New Kuala Lumpur, Malaysia

Bybit Limited • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Study Growth Fund
Internal Events
Global Collaboration
+2
Principal Tech & Product Engineer
Principal Tech & Product Engineer

Bybit • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Study Growth Fund
Internal Events
Global Collaboration
+2
SRE Leader Kuala Lumpur, Malaysia
SRE Leader Kuala Lumpur, Malaysia

Bybit Limited • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Risk Control Data Analysis & Operations Engineer
Risk Control Data Analysis & Operations Engineer

Bybit • Kuala Lumpur

On-site
MYR 60,000 - 100,000
Study Growth Fund
Internal Events
Global Collaboration
+2
Devops Expert Kuala Lumpur, Malaysia
Devops Expert Kuala Lumpur, Malaysia

Bybit Limited • Kuala Lumpur

On-site
MYR 300,000 - 520,000
Study Growth Fund
Internal events
Global collaboration
+2
(Principal) Golang Development Engineer
(Principal) Golang Development Engineer

Bybit • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Study Growth Fund
Internal Events
Global Collaboration
+2
Principal Backend Development Engineer New Kuala Lumpur, Malaysia
Principal Backend Development Engineer New Kuala Lumpur, Malaysia

Bybit Limited • Kuala Lumpur

On-site
MYR 120,000 - 240,000
Study Growth Fund
Internal Events
Global Collaboration
+2
Lead Security Management Engineer
Lead Security Management Engineer

Bybit • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Growth fund
Internal events
Global collaboration
+2
Lead Frontend Development Engineer
Lead Frontend Development Engineer

Bybit • Kuala Lumpur

On-site
MYR 90,000 - 150,000
Study Growth Fund
Internal Events
Global Collaboration
+2
[GRC] Risk Control Data Analysis & Operations Engineer
[GRC] Risk Control Data Analysis & Operations Engineer

Unchain Data • Kuala Lumpur

On-site
Study Growth Fund
Internal Events
Global Collaboration
+2