Cloud Site Relibility Engineer - DCS

BYTEDANCE PTE. LTD.

Singapore

On-site

SGD 100,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ByteDance is seeking a highly skilled DevOps/SRE engineer to design, build, and operate its global infrastructure. You’ll join the DCS team responsible for hyperscale datacenters, cloud platforms, and reliable services across public and private environments.

You will develop automation, monitoring, and tooling, manage image/AMI pipelines, and participate in on-call rotations to address incidents related to cloud, OS, network, and performance. Strong collaboration is essential.

Qualifications

  • Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • 2+ years of experience in Linux operations, SRE, or DevOps; experience operating large-scale production environments is a strong plus.
  • Proficient in Go, Python, or C++, with solid engineering capabilities in platform development, system tooling, and automation.
  • Strong CS fundamentals, with deep understanding of Linux OS principles, networks, storage, GPUs, and databases, plus troubleshooting and RCA skills.
  • Familiar with reliability practices: monitoring/alerting, capacity management, change management, canary/gray releases, incident response, postmortem.

Responsibilities

  • Design, build, scale, and operate ByteDance’s global infrastructure, including large-scale systems spanning public and private clouds.
  • Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure.
  • Create, manage, and standardize cloud AMIs/images for use across environments with strict compliance alignment.
  • Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability.
  • Drive improvements across the infrastructure lifecycle from ideation to deployment, usage support, and ongoing refinement.

Skills

Linux operations
SRE / DevOps
Go / Python / C++
Automation & tooling
Networking & storage
Monitoring & incident response
Cloud & virtualization
Communication

Education

Bachelor’s degree or above in CS/SE/InfoSec or related field

Tools

Docker
Kubernetes
KVM

Job description

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Responsibilities

About the TeamThe DCS team supports the company's fast growth by building and operating hyperscale datacenters. The team manages the end to end lifecycle of server fleet, providing cloud solutions and various infrastructure services ensuring that they are scalable and are reliable.

  • Design, build, scale, and operate ByteDance’s global infrastructure, including large-scale systems spanning public and private clouds.
  • Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure.
  • Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the company's global compliance standards.
  • Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability.
  • Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.
Qualifications
Minimum Qualifications:
  • Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • 2+ years of experience in Linux operations, SRE, or DevOps; experience operating large-scale production environments is a strong plus.
  • Proficient in at least one programming language such as Go, Python, or C++, with solid engineering capabilities in platform development, system tooling, and automation.
  • Strong computer science fundamentals, with deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, along with systematic troubleshooting and root-cause analysis skills.
  • Familiar with core reliability practices, including monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and postmortem processes.
  • Strong communication and collaboration skills, with the ability to proactively identify problems, drive cross-team execution, and demonstrate strong ownership and results-oriented mindset.
Preferred Qualifications:
  • Hands-on experience operating public cloud platforms, or deep familiarity with major cloud providers such as OCI, AWS, Azure, GCP, etc, including understanding of their underlying mechanisms.
  • Experience with large-scale cloud host delivery, image/AMI systems, resource scheduling, network adaptation, and virtualization technologies such as KVM/QEMU.
  • Familiar with containers and cloud-native ecosystems, including Docker, Kubernetes, and containerd, with a solid understanding of isolation mechanisms like cgroups and namespaces.
  • Experience maintaining GPU clusters, including drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, and GPU delivery pipelines.
  • Proven experience in reliability-focused initiatives such as failure drill systems, capacity governance, change governance, observability platforms, and resource cost optimization.
  • Open-source contributions, technical blogs, patents, or technical sharing experience are highly preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Site Relibility Engineer - DCS Singapore Regular
Cloud Site Relibility Engineer - DCS Singapore Regular

ByteDance • Singapore

On-site
SGD 75,000 - 100,000
Data Center Infrastructure Global FOC Shift Supervisor, DCS
Data Center Infrastructure Global FOC Shift Supervisor, DCS

ByteDance • Singapore

On-site
SGD 60,000 - 120,000
Backend Software Engineer (SRE) - Cloud Infrastructure
Backend Software Engineer (SRE) - Cloud Infrastructure

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Global Edge Delivery and Operation Lead, DCS
Global Edge Delivery and Operation Lead, DCS

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 180,000 - 260,000
Meals provided
Competitive compensation
Site Reliability Engineer (Cloud) - Infrastructure Engineering Technology - DevOps Singapore Regular
Site Reliability Engineer (Cloud) - Infrastructure Engineering Technology - DevOps Singapore Regular

ByteDance • Singapore

On-site
SGD 60,000 - 90,000
Site Reliability Engineer (Cloud) - Infrastructure Engineering
Site Reliability Engineer (Cloud) - Infrastructure Engineering

ByteDance • Singapore

On-site
SGD 180,000 - 240,000
Data Center Infrastructure Global FOC Shift Supervisor, DCS
Data Center Infrastructure Global FOC Shift Supervisor, DCS

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer, System - System Service Global
Site Reliability Engineer, System - System Service Global

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Global Edge Delivery and Operation Lead, DCS
Global Edge Delivery and Operation Lead, DCS

ByteDance • Singapore

On-site
SGD 180,000 - 320,000
Global Data Center Infrastructure Technology Expert, DCS
Global Data Center Infrastructure Technology Expert, DCS

ByteDance • Singapore

On-site
SGD 70,000 - 110,000