Senior Principal Site Reliability Engineer

Bybit

Hong Kong

On-site

HKD 1,200,000 - 2,000,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Bybit is seeking a Senior Chaos/Resilience Engineer to design and operate an enterprise-grade chaos engineering platform across multi-cluster, multi-region deployments. You will own fault-injection tooling, safety controls, and automated recovery, partnering with SRE and platform teams to raise resilience across the production stack.

You will mentor 2–3 engineers, stay ahead of industry trends, and define playbooks for incident injection, monitoring, and SLO alignment.

Qualifications

  • 8+ years backend/infrastructure engineering experience, with 3+ years in chaos or stability engineering.
  • Hands-on experience with large-scale fault injection in production environments.
  • Expert-level Kubernetes fault injection (Chaos Mesh / Litmus / custom solutions), familiar with CRD/Operator development.
  • Proficient in Go (or other backend language) with architecture design capability.
  • Deep understanding of distributed system failure modes.
  • Familiarity with observability stacks (Prometheus / Grafana / OpenTelemetry).
  • Excellent technical documentation and solution design skills.

Responsibilities

  • Design and build an enterprise-grade chaos engineering platform across multi-cluster, multi-region deployments.
  • Develop fault injection engine and safety controls for production environments.
  • Production safety: blast radius control, one-click Kill Switch, automated rollback, real-time impact monitoring.
  • Integrate with monitoring and SLO systems to automate fault observation and pass/fail decisions.
  • Mentor 2–3 engineers in chaos engineering and stay current with industry practices (AI-driven fault scenarios).
  • Collaborate with SREs and developers to improve system resilience across services.

Skills

Kubernetes
Chaos engineering
Go language
Observability
Documentation
Leadership

Tools

Chaos Mesh
Litmus
OpenTelemetry
Prometheus
Grafana
AWS

Job description

About Us

Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance.

Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution.

Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services.

Core Responsibilities
Chaos Engineering Platform Architecture & Development (50%)
  • Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (K8s + EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments
Core capability development
  • Fault Injection Engine: Pod-level / Node-level / AZ-level fault simulation, network latency / packet loss / partition, dependency timeout / error injection
  • Production Safety Assurance: Blast radius control, one-click Kill Switch, automatic rollback, real-time impact monitoring
  • Traffic Isolation: Experiment traffic tagging and isolation to ensure fault injection does not impact real users
  • Fault Isolation: Precise impact scoping at service / cluster / AZ granularity
  • Design experiment orchestration capabilities supporting complex fault scenario composition (e.g., simultaneous network latency + downstream timeout + cache invalidation)
  • Deep integration with existing monitoring, alerting, and SLO systems to achieve an automated closed loop: inject fault → observe impact → determine pass/fail
Production Resilience Validation Framework (30%)
  • Define safety standards and approval workflows for mainnet fault injection
  • Design and drive routine chaos experiments:
  • Daily patrol-level experiments: Low-risk experiments executed automatically on a daily/weekly basis
  • Periodic validation experiments: Monthly/quarterly resilience verification of critical paths
  • Large-scale drills: Cross-AZ / cross-region disaster recovery failover validation
  • Establish a resilience scoring system to quantify system health based on experiment results
  • Deliver improvement recommendations and drive business teams to remediate identified weaknesses
3. Technology Selection & Team Enablement (20%)
  • Evaluate and select the technology foundation (Chaos Mesh / Litmus / custom components — hybrid strategy)
  • Develop chaos engineering best practices and playbooks to enable SRE teams and application developers
  • Mentor and grow the team (2–3 engineers) in chaos engineering capabilities
  • Stay current with industry developments and introduce cutting-edge practices (e.g., AI-driven fault scenario discovery)
Requirements
Must-Have:
  • 8+ years of backend / infrastructure engineering experience, with 3+ years dedicated to chaos engineering or stability engineering
  • Hands-on experience with large-scale fault injection in production environments (not just test environments), with deep understanding of production safety constraints
  • Expert-level proficiency in Kubernetes fault injection (Chaos Mesh / Litmus / custom solutions), familiar with CRD / Operator development
  • Proficient in at least one backend language (Go preferred), with platform-level system architecture design capability
  • Deep understanding of distributed system failure modes (network partitions, split-brain, cascading failures, data inconsistency, etc.)
  • Familiarity with observability tech stack (Prometheus / Grafana / Thanos / OpenTelemetry)
  • Excellent technical documentation and solution design skills
Nice-to-Have:
  • Experience in financial / trading system stability (understanding of transaction consistency and fund safety constraints)
  • Experience building SLO / Error Budget frameworks
  • Experience building automated fault recovery (self-healing) systems
  • Familiarity with AWS infrastructure (EC2 / EKS / Multi-AZ / Multi-Region)
  • Knowledge of Netflix Chaos Engineering / AWS Fault Injection Simulator / Gremlin
  • Open-source community contributions (Chaos Mesh / Litmus or similar projects)
Soft Skills:
  • Ability to balance "safety" and "validation depth" — not afraid of production injection, while maintaining strict risk control
  • Strong cross-team collaboration and influence — chaos engineering requires buy-in from business teams; this role demands persuasion skills
  • Self-driven, capable of independently planning and executing in ambiguous situations
Why Join Us

At Bybit, we are committed to fostering a supportive and enriching work environment.

Our benefits include:

  • Study Growth Fund: We support your professional development and continuous learning.
  • Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation.
  • Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world.
  • Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company.
  • Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Principal Chaos Engineer, Platform & Resilience
Senior Principal Chaos Engineer, Platform & Resilience

Bybit • Hong Kong

On-site
HKD 1,200,000 - 2,000,000
Principal Security Operation Engineer (Red Team) New Hong Kong SAR; Kuala Lumpur, Malaysia
Principal Security Operation Engineer (Red Team) New Hong Kong SAR; Kuala Lumpur, Malaysia

Bybit Limited • Hong Kong

On-site
HKD 900,000 - 1,500,000
Study Growth Fund
Internal Events
Global Collaboration
+2
Principal Rust Development Engineer
Principal Rust Development Engineer

Bybit • Hong Kong

On-site
HKD 900,000 - 1,200,000
Growth Fund
Internal events
Global collaboration
+2
Senior Engineer, Risk Engineering
Senior Engineer, Risk Engineering

P2P • Hong Kong

On-site
HKD 600,000 - 1,000,000
Senior Admin Specialist (Contractor)
Senior Admin Specialist (Contractor)

Bybit • Hong Kong

On-site
HKD 167,000 - 257,000
Study Growth Fund
Internal Events
Global Collaboration
+2
Java Development Engineer Intern (OTC Team)
Java Development Engineer Intern (OTC Team)

Bybit • Hong Kong

On-site
HKD 520,000 - 1,000,000
Senior Brokerage Systems Engineer (Trading, Clearing & Settlement, Accounting Systems)
Senior Brokerage Systems Engineer (Trading, Clearing & Settlement, Accounting Systems)

BIT Official • Hong Kong

On-site
HKD 1,000,000 - 1,500,000
DevOps Engineer
DevOps Engineer

BIT Official • Hong Kong

On-site
HKD 600,000 - 1,000,000
Compliance & AML Lead, Hong Kong
Compliance & AML Lead, Hong Kong

Bybit • Hong Kong

On-site
HKD 1,200,000 - 1,800,000
Study Growth Fund
Internal Events
Global Collaboration
+2
Senior Rust Engineer (Performance engineering)
Senior Rust Engineer (Performance engineering)

Binance • Hong Kong

Hybrid
HKD 600,000 - 900,000
Work-from-home arrangement
Competitive salary