Global SRE: Core Data Center & Service Reliability

ByteDance

San Jose (CA)

On-site

USD 162,000 - 388,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

ByteDance is seeking a self-motivated system engineer with an SRE mindset to manage large-scale Linux host infrastructure across non-China data centers. You will own deployment architectures, fleet health, and core services like DNS, NTP, DHCP, and Kerberos.

Collaborate with network and security teams to meet evolving business demands while driving reliability improvements through SLOs, incident response, and post-mortems.

Qualifications

  • Bachelor’s degree or higher in Electrical Engineering, Computer Engineering, Computer Science or related majors.
  • Solid experience managing large-scale Linux hosts, OS deployment, configuration management, patching, and fleet operations.
  • Hands-on knowledge of core data center services: DNS (BIND/PowerDNS), NTP, DHCP, NAT, APT repository management, and Kerberos.
  • Proficiency with DevOps tooling and configuration management tools (e.g., Ansible, Salt, Puppet) and CI/CD pipelines.
  • Familiarity with SRE principles, SLO/SLI, error budgets, and blameless post-mortems.
  • Understanding of high-availability design patterns and disaster recovery.

Responsibilities

  • Manage and maintain large-scale host infrastructure across ByteDance's non-China data centers, including OS lifecycle management, configuration standardization, and fleet health monitoring.
  • Own reliability and availability of core data center services such as DNS, NTP, DHCP, NAT, APT repository, and Kerberos authentication.
  • Design and implement deployment architectures for foundational services with high availability, fault tolerance, and multi-region disaster recovery.
  • Develop and enforce SLOs for managed services; lead incident response, root cause analysis, and post-mortems to drive reliability improvements.
  • Collaborate with network, security, and application teams to meet evolving business demands.
  • Identify automation opportunities across host management and service operations; drive tooling and process improvements to reduce toil and increase efficiency.

Skills

SRE mindset
DevOps
Linux administration
Incident response
Automation

Education

Bachelor’s degree or higher in Electrical Engineering, Computer Engineering, Computer Science or related majors

Tools

Ansible
Salt
Puppet
CI/CD pipelines
DNS
NTP
Kerberos

Job description

ByteDance is seeking a self-motivated system engineer with an SRE mindset to manage large-scale Linux host infrastructure across non-China data centers. You will own deployment architectures, fleet health, and core services like DNS, NTP, DHCP, and Kerberos.

Collaborate with network and security teams to meet evolving business demands while driving reliability improvements through SLOs, incident response, and post-mortems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Graduate: Data Platform Reliability & Automation
SRE Graduate: Data Platform Reliability & Automation

ByteDance • San Jose (CA)

On-site
USD 150,000 - 190,000
Senior SRE - Data Infrastructure & Reliability
Senior SRE - Data Infrastructure & Reliability

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Senior SRE: Scale, Reliability & Incident Mastery
Senior SRE: Scale, Reliability & Incident Mastery

ByteDance • Seattle (WA)

On-site
USD 212,000 - 368,000
Cloud SRE Tech Lead - Global Infra & Automation
Cloud SRE Tech Lead - Global Infra & Automation

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Senior Global Server Ops & Reliability Engineer
Senior Global Server Ops & Reliability Engineer

Bytedance • San Jose (CA)

On-site
USD 122,000 - 272,000
Tech Lead, Data Infrastructure & SRE (Cloud-Scale)
Tech Lead, Data Infrastructure & SRE (Cloud-Scale)

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Health insurance
401(k) with company match
Paid parental leave
+3
Tech Lead: Data Infrastructure & Site Reliability
Tech Lead: Data Infrastructure & Site Reliability

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500
Medical insurance
Dental and vision insurance
401(k) with company match
+7
Senior Server Reliability & NPI Operations Engineer
Senior Server Reliability & NPI Operations Engineer

ByteDance • San Jose (CA)

On-site
USD 115,200 - 288,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Global Data Center Operations Supervisor
Global Data Center Operations Supervisor

Pangle • Ashburn (VA)

Hybrid
USD 86,000 - 167,000
Medical, dental & vision
401(k) with company match
Paid parental leave
+2
Global SRE: Multi-Region Infra & Reliability Engineer
Global SRE: Multi-Region Infra & Reliability Engineer

rednote • Palo Alto (CA)

On-site
USD 200,000 - 400,000