Site Reliability Engineer, System - System Service Global

ByteDance

Singapore

On-site

SGD 90,000 - 150,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ByteDance’s Global System Service team operates the data-center infrastructure powering ByteDance’s non-China regions. We seek a self-motivated system engineer with SRE and DevOps skills to manage large-scale Linux host infrastructure, ensure core services reliability, and design HA architectures across regions.

You will develop automation, implement deployment architectures, and drive incident response with blameless post-mortems, collaborating with network, security, and application teams to

Qualifications

  • Bachelor's degree in Electrical Engineering, Computer Engineering, or Computer Science
  • Experience in large-scale Linux host management including OS deployment, configuration management, patching, and fleet operations
  • Proficiency with core data center services: DNS (BIND/PowerDNS), NTP, DHCP, NAT, APT repo management, and Kerberos
  • Experience with configuration management tools (Ansible, Salt, Puppet) and CI/CD pipelines
  • Familiarity with SRE principles including SLO/SLI, error budgets, and blameless post‑mortems
  • Understanding of high availability designs and disaster recovery strategies
  • Strong troubleshooting across Linux systems and networking

Responsibilities

  • Manage and maintain large-scale host infrastructure across ByteDance's non-China data centers
  • Own reliability and availability of core data center services (DNS, NTP, DHCP, NAT, APT, Kerberos)
  • Design and implement deployment architectures for foundational services with high availability and DR
  • Develop and enforce SLOs; lead incident response and post-mortem reviews
  • Collaborate with network, security, and application teams to meet global growth needs
  • Identify automation opportunities; drive tooling to reduce toil and increase efficiency

Skills

Linux
SRE
DevOps
Ansible
Puppet
Salt
CI/CD
DNS

Education

Bachelor's degree in Electrical/Computer Engineering or Computer Science

Tools

DNS (BIND/PowerDNS)
NTP
DHCP
Kerberos
APT repo management

Job description

Responsibilities

About the Team

The Global System Service team owns the infrastructure services and management solutions that power ByteDance's data centers outside of China — from day-to-day operations to long-term architecture design and maintenance. The team specializes in composing end-to-end solutions by drawing on both open-source community tools and in-house developed products, tailored to both the business requirements and the operational complexities of large-scale infrastructure across ByteDance's non-China regions. Our mission is to deliver efficient infrastructure solutions and a stable, secure system environment for ByteDance's global business.

We are looking for a self‑motivated system engineer that is equipped with SRE mindset and DevOps skills. Your responsibilities will include:

  • Manage and maintain large-scale host infrastructure across ByteDance's non-China data centers, covering OS lifecycle management, configuration standardization, and fleet-wide health monitoring.
  • Own the reliability and availability of core data center foundational services, including DNS, NTP, DHCP, NAT, APT repository, and Kerberos authentication.
  • Design and implement deployment architectures for foundational services, ensuring high availability, fault tolerance, and disaster recovery across regions.
  • Develop and enforce SLOs for managed services; lead incident response, root cause analysis, and post‑mortem reviews to drive continuous reliability improvements.
  • Collaborate with network, security, and application teams to ensure foundational services meet the evolving demands of global business growth.
  • Identify automation opportunities across host management and service operations; drive tooling and process improvements to reduce toil and increase operational efficiency.
Qualifications
Minimum Qualifications
  • Bachelor’s degree or higher in Electrical Engineering, Computer Engineering, Computer Science or related majors.
  • Solid experience in large-scale Linux host management, including OS deployment, configuration management, patching, and fleet operations.
  • Strong hands‑on knowledge of core data center foundational services: DNS (BIND/PowerDNS), NTP, DHCP, NAT, APT repository management, and Kerberos.
  • Proficiency with DevOps tooling, including configuration management tools (e.g., Ansible, Salt, Puppet) and CI/CD pipelines.
  • Familiarity with SRE principles and practices, including SLO/SLI definition, error budget management, and blameless post‑mortems.
  • Solid understanding of high availability design patterns, active‑active/active‑passive architectures, and disaster recovery strategies.
  • Strong troubleshooting skills across the Linux system stack and network layer.
Preferred Qualifications
  • Experience managing host fleets at scale (thousands of nodes or above) in a production environment.
  • Scripting or development experience in Python, Go, or Bash for automation and tooling.
  • Exposure to hybrid or multi-region data center environments.
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, System - System Service Global
Site Reliability Engineer, System - System Service Global

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer (Cloud) - Infrastructure Engineering
Site Reliability Engineer (Cloud) - Infrastructure Engineering

ByteDance • Singapore

On-site
SGD 180,000 - 240,000
Software Engineer (SRE - Platform Services), Infrastructure Engineering
Software Engineer (SRE - Platform Services), Infrastructure Engineering

Bytedance • Singapore

On-site
SGD 100,000 - 150,000
Tech Lead (SRE) - Cloud Infrastructure
Tech Lead (SRE) - Cloud Infrastructure

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 180,000 - 240,000
Backend Software Engineer (SRE) - Cloud Infrastructure
Backend Software Engineer (SRE) - Cloud Infrastructure

ByteDance • Singapore

On-site
SGD 140,000 - 210,000
Software Engineer (SRE - Platform Services), Infrastructure Engineering
Software Engineer (SRE - Platform Services), Infrastructure Engineering

re-zoo-me • Singapore

Hybrid
SGD 90,000 - 130,000
Site Reliability Engineer - Big Data Computer Platform
Site Reliability Engineer - Big Data Computer Platform

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Cloud Site Relibility Engineer - DCS
Cloud Site Relibility Engineer - DCS

ByteDance • Singapore

On-site
SGD 110,000 - 180,000
Software Engineer (SRE - Platform Services), Infrastructure Engineering
Software Engineer (SRE - Platform Services), Infrastructure Engineering

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 90,000 - 140,000
Site Reliability Engineer - Singapore
Site Reliability Engineer - Singapore

ByteDance • Singapore

On-site
SGD 100,000 - 180,000