Global Data Center SRE: Reliability & Automation

BYTEDANCE PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ByteDance's Global System Service team powers data center infrastructure outside China, managing OS deployment, fleet health, and centralized configuration across multiple regions. We seek a self-motivated system engineer with an SRE mindset to own reliability of DNS, NTP, DHCP, NAT, and Kerberos, design HA deployments, and lead post-mortems to drive continuous improvement.

Join us to partner with network, security, and application teams, automate host management, and reduce toil at scale in a

Qualifications

  • Bachelor's degree or higher in Electrical Engineering, Computer Engineering, Computer Science or related majors.
  • Experience in large-scale Linux host management and OS lifecycle.
  • Strong knowledge of core data center services: DNS, NTP, DHCP, NAT, APT repo, Kerberos.
  • Proficiency with DevOps tools (Ansible, Salt, Puppet) and CI/CD pipelines.
  • Familiarity with SRE principles, including SLO/SLI, error budgets, and post-mortems.
  • Understanding of HA design patterns and disaster recovery.
  • Troubleshooting across Linux and network layers.

Responsibilities

  • Manage and maintain large-scale host infrastructure across ByteDance's non-China data centers, covering OS lifecycle management, configuration standardization, and fleet-wide health monitoring.
  • Own the reliability and availability of core data center foundational services, including DNS, NTP, DHCP, NAT, APT repository, and Kerberos authentication.
  • Design and implement deployment architectures for foundational services, ensuring high availability, fault tolerance, and disaster recovery across regions.
  • Develop and enforce SLOs for managed services; lead incident response, root cause analysis, and post-mortem reviews to drive continuous reliability improvements.
  • Collaborate with network, security, and application teams to ensure foundational services meet the evolving demands of global business growth.
  • Identify automation opportunities across host management and service operations; drive tooling and process improvements to reduce toil and increase operational efficiency.

Skills

Linux host management
SRE principles
DevOps tooling
DNS/NTP/DHCP/NAT
Scripting (Python/Go/Bash)
Troubleshooting

Education

Bachelor's degree in Electrical Engineering/CS or related

Tools

Ansible
Salt
Puppet
CI/CD pipelines
Python

Job description

ByteDance's Global System Service team powers data center infrastructure outside China, managing OS deployment, fleet health, and centralized configuration across multiple regions. We seek a self-motivated system engineer with an SRE mindset to own reliability of DNS, NTP, DHCP, NAT, and Kerberos, design HA deployments, and lead post-mortems to drive continuous improvement.

Join us to partner with network, security, and application teams, automate host management, and reduce toil at scale in a

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Global SRE Engineer: Large-Scale Infra & Reliability
Global SRE Engineer: Large-Scale Infra & Reliability

ByteDance • Singapore

On-site
SGD 150,000 - 190,000
Site Reliability Engineer, System - System Service Global
Site Reliability Engineer, System - System Service Global

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Cloud SRE: Scale Global Infra & Automation
Cloud SRE: Scale Global Infra & Automation

ByteDance • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer, System - System Service Global
Site Reliability Engineer, System - System Service Global

ByteDance • Singapore

On-site
SGD 150,000 - 190,000
Cloud Infrastructure Backend Engineer | SRE & Automation
Cloud Infrastructure Backend Engineer | SRE & Automation

ByteDance • Singapore

On-site
SGD 120,000 - 180,000
Global Data Center Infrastructure Engineer
Global Data Center Infrastructure Engineer

Pangleglobal • Singapore

On-site
SGD 70,000 - 100,000
SRE Tech Lead — Cloud Infrastructure & Reliability
SRE Tech Lead — Cloud Infrastructure & Reliability

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 180,000 - 240,000
Cloud Infra SRE Engineer: Reliability & Automation
Cloud Infra SRE Engineer: Reliability & Automation

ByteDance • Singapore

On-site
SGD 55,000 - 90,000
Global Data Center Infra O&M Shift Supervisor
Global Data Center Infra O&M Shift Supervisor

ByteDance • Singapore

On-site
SGD 60,000 - 90,000
Hybrid Cloud SRE: Delivery & Reliability Engineer
Hybrid Cloud SRE: Delivery & Reliability Engineer

ByteDance • Singapore

Hybrid
SGD 80,000 - 120,000