Site Reliability Engineer, System - System Service Global

ByteDance

Singapore

On-site

SGD 150,000 - 190,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a systems engineer with a strong SRE mindset to manage and scale Linux host infrastructure across non-China data centers. You will own DNS, NTP, DHCP, NAT, and Kerberos services and design deployment architectures for high availability and disaster recovery.

The role requires hands-on experience with Ansible/Salt/Puppet, CI/CD pipelines, and mature incident response practices. Collaboration with cross-functional teams is essential to support global operations.

Qualifications

  • Bachelor's degree or higher in Electrical/Computer Engineering, Computer Engineering, Computer Science or related majors.
  • Solid experience in large-scale Linux host management, including OS deployment and fleet operations.
  • Strong hands-on knowledge of core data center services: DNS, NTP, DHCP, NAT, APT repository management, Kerberos.
  • Proficiency with DevOps tooling and CI/CD pipelines.
  • Familiarity with SRE principles, including SLO/SLI and post-mortems.
  • Understanding of high availability design patterns and disaster recovery.

Responsibilities

  • Manage and maintain large-scale host infrastructure across ByteDance's non-China data centers.
  • Own reliability and availability of core data center services (DNS, NTP, DHCP, NAT, APT, Kerberos).
  • Design and implement deployment architectures for foundational services with high availability and DR across regions.
  • Develop and enforce SLOs; lead incident response and post-mortems to improve reliability.
  • Collaborate with network, security, and application teams to meet global demand.
  • Identify automation opportunities to reduce toil and improve efficiency.

Skills

Linux system management
DNS/NTP/DHCP/NAT
SRE principles
DevOps tooling (Ansible,Salt,Puppet)
CI/CD pipelines
Linux and network troubleshooting

Education

Bachelor's degree in Electrical/Computer Engineering or Computer Science

Tools

Ansible
Salt
Puppet
Python/Go scripting

Job description

Responsibilities
About the Team

The Global System Service team owns the infrastructure services and management solutions that power ByteDance's data centers outside of China — from day-to-day operations to long-term architecture design and maintenance. The team specializes in composing end-to-end solutions by drawing on both open-source community tools and in-house developed products, tailored to both the business requirements and the operational complexities of large-scale infrastructure across ByteDance's non-China regions. Our mission is to deliver efficient infrastructure solutions and a stable, secure system environment for ByteDance's global business.

Responsibilities

We are looking for a self-motivated system engineer that is equipped with SRE mindset and DevOps skills. Your responsibilities will include:

  • Manage and maintain large-scale host infrastructure across ByteDance's non-China data centers, covering OS lifecycle management, configuration standardization, and fleet-wide health monitoring.
  • Own the reliability and availability of core data center foundational services, including DNS, NTP, DHCP, NAT, APT repository, and Kerberos authentication.
  • Design and implement deployment architectures for foundational services, ensuring high availability, fault tolerance, and disaster recovery across regions.
  • Develop and enforce SLOs for managed services; lead incident response, root cause analysis, and post-mortem reviews to drive continuous reliability improvements.
  • Collaborate with network, security, and application teams to ensure foundational services meet the evolving demands of global business growth.
  • Identify automation opportunities across host management and service operations; drive tooling and process improvements to reduce toil and increase operational efficiency.
Qualifications

Minimum Qualifications:

  • Bachelor’s degree or higher in Electrical Engineering, Computer Engineering, Computer Science or related majors.
  • Solid experience in large-scale Linux host management, including OS deployment, configuration management, patching, and fleet operations.
  • Strong hands-on knowledge of core data center foundational services: DNS (BIND/PowerDNS), NTP, DHCP, NAT, APT repository management, and Kerberos.
  • Proficiency with DevOps tooling, including configuration management tools (e.g., Ansible, Salt, Puppet) and CI/CD pipelines.
  • Familiarity with SRE principles and practices, including SLO/SLI definition, error budget management, and blameless post-mortems.
  • Solid understanding of high availability design patterns, active-active/active-passive architectures, and disaster recovery strategies.
  • Strong troubleshooting skills across the Linux system stack and network layer.

Preferred Qualifications:

  • Experience managing host fleets at scale (thousands of nodes or above) in a production environment.
  • Scripting or development experience in Python, Go, or Bash for automation and tooling.
  • Exposure to hybrid or multi-region data center environments.
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend Software Engineer (SRE) - Cloud Infrastructure
Backend Software Engineer (SRE) - Cloud Infrastructure

ByteDance • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer - Traffic Infrastructure
Site Reliability Engineer - Traffic Infrastructure

ByteDance • Singapore

On-site
SGD 90,000 - 150,000
Backend Software Engineer (SRE) Graduate (Cloud Infrastructure) - 2027 Start
Backend Software Engineer (SRE) Graduate (Cloud Infrastructure) - 2027 Start

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 67,000 - 112,000
Site Reliability Engineer (Cloud) - Infrastructure Engineering Technology - DevOps Singapore Regular
Site Reliability Engineer (Cloud) - Infrastructure Engineering Technology - DevOps Singapore Regular

ByteDance • Singapore

On-site
SGD 60,000 - 90,000
Site Reliability Engineer System Graduate (System Service Global (Infrastructure Engineering)) [...]
Site Reliability Engineer System Graduate (System Service Global (Infrastructure Engineering)) [...]

ByteDance • Singapore

On-site
SGD 60,000 - 90,000
Backend Software Engineer (SRE) Graduate (Cloud Infrastructure) - 2027 Start
Backend Software Engineer (SRE) Graduate (Cloud Infrastructure) - 2027 Start

ByteDance • Singapore

On-site
SGD 55,000 - 90,000
Site Reliability Engineer Intern, System - System Service Global (Infrastructure Engineering), [...]
Site Reliability Engineer Intern, System - System Service Global (Infrastructure Engineering), [...]

ByteDance • Singapore

On-site
SGD 13,000 - 27,000
Cloud Site Relibility Engineer - DCS
Cloud Site Relibility Engineer - DCS

ByteDance • Singapore

On-site
SGD 90,000 - 150,000
Big Data SRE Operations Expert - Computer Platform Singapore Regular
Big Data SRE Operations Expert - Computer Platform Singapore Regular

ByteDance • Singapore

On-site
SGD 70,000 - 90,000
Site Reliability Engineer Intern (Video and Edge, CDN Platform) - 2027 Start
Site Reliability Engineer Intern (Video and Edge, CDN Platform) - 2027 Start

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 17,000 - 30,000