Cloud Site Relibility Engineer - DCS Singapore Regular

ByteDance

Singapore

On-site

SGD 75,000 - 100,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

ByteDance in Singapore is looking for a professional to join their team in managing and optimizing global infrastructure. This position emphasizes building scalable systems across public and private clouds, significantly improving operations through tools and automation frameworks. Candidates should have a Bachelor's degree in Computer Science or a similar field, along with at least 2 years of relevant experience in Linux operations or DevOps. Knowledge of cloud providers such as AWS, Azure, or GCP is essential.

Qualifications

  • 2+ years of experience in Linux operations, SRE, or DevOps.
  • Strong communication and collaboration skills.
  • Proficient in at least one programming language.

Responsibilities

  • Design, build, scale, and operate ByteDance’s global infrastructure.
  • Develop tools, automation frameworks, visualizations, and monitoring systems.
  • Drive improvements across the entire infrastructure lifecycle.

Skills

Linux operations
SRE
DevOps
Go
Python
C++

Education

Bachelor’s degree in Computer Science or related field

Tools

Docker
Kubernetes
AWS
Azure
GCP

Job description

About the Team

The DCS team supports the company's fast growth by building and operating hyperscale datacenters. The team manages the end‑to‑end lifecycle of the server fleet, providing cloud solutions and various infrastructure services to ensure they are scalable and reliable.

Responsibilities
  • Design, build, scale, and operate ByteDance’s global infrastructure, including large‑scale systems spanning public and private clouds.
  • Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure.
  • Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the company's global compliance standards.
  • Thrive in a fast‑paced environment, engaging in technical operations and on‑call rotations to address incidents related to cloud, OS, network, performance, and reliability.
  • Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.
Qualifications
Minimum Qualifications
  • Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • 2+ years of experience in Linux operations, SRE, or DevOps.
  • Proficient in at least one programming language such as Go, Python, or C++, with solid engineering capabilities in platform development, system tooling, and automation.
  • Strong computer science fundamentals, with deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, along with systematic troubleshooting and root‑cause analysis skills.
  • Familiar with core reliability practices, including monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and post‑mortem processes.
  • Strong communication and collaboration skills, with the ability to proactively identify problems, drive cross‑team execution, and demonstrate strong ownership and results‑oriented mindset.
Preferred Qualifications
  • Hands‑on experience operating public cloud platforms, or deep familiarity with major cloud providers such as OCI, AWS, Azure, GCP, etc., including understanding of their underlying mechanisms.
  • Experience with large‑scale cloud host delivery, image/AMI systems, resource scheduling, network adaptation, and virtualization technologies such as KVM/QEMU.
  • Familiar with containers and cloud‑native ecosystems, including Docker, Kubernetes, and container environments, with a solid understanding of isolation mechanisms like cgroups and namespaces.
  • Experience maintaining GPU clusters, including drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, and GPU delivery pipelines.
  • Proven experience in reliability‑focused initiatives such as failure drill systems, capacity governance, change governance, observability platforms, and resource cost optimization.
  • Open‑source contributions, technical blogs, patents, or technical sharing experience are highly preferred.
  • Experience operating large‑scale production environments is a strong plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid Cloud Operation and Delivery Engineer (SRE) - Data Infrastructure Singapore Regular
Hybrid Cloud Operation and Delivery Engineer (SRE) - Data Infrastructure Singapore Regular

ByteDance • Singapore

Hybrid
SGD 80,000 - 120,000
Cloud Site Reliability Engineer
Cloud Site Reliability Engineer

SINGAPORE POOLS (PRIVATE) LIMITED. • Singapore

On-site
SGD 120,000 - 160,000
Total rewards
Health benefits
Learning opportunities
+1
Senior Cloud SRE: Global Infrastructure & Automation
Senior Cloud SRE: Global Infrastructure & Automation

ByteDance • Singapore

On-site
SGD 180,000 - 240,000
Data Center Facilities Manager, DCS Singapore Regular
Data Center Facilities Manager, DCS Singapore Regular

Pangleglobal • Singapore

On-site
SGD 90,000 - 120,000
Cloud Site Relibility Engineer Intern (DCS), 2027 Start
Cloud Site Relibility Engineer Intern (DCS), 2027 Start

ByteDance • Singapore

On-site
SGD 11,000 - 22,000
Global Cloud SRE: Reliability, Scale & Automation
Global Cloud SRE: Reliability, Scale & Automation

ByteDance • Singapore

On-site
SGD 75,000 - 100,000
Cloud Infra Engineer
Cloud Infra Engineer

Centre for Strategic Infocomm Technologies (CSIT) • Singapore

On-site
SGD 60,000 - 80,000
Cloud Engineer
Cloud Engineer

Envoy Search Partners Pte Limited • Singapore

On-site
SGD 90,000 - 130,000
Flexi benefits covering medical, gym,
Cloud SRE Intern: Build Reliable Big Data Platforms
Cloud SRE Intern: Build Reliable Big Data Platforms

Tencent • Singapore

On-site
SGD 60,000 - 100,000
Cloud Engineer
Cloud Engineer

Tesla Consulting Group • Singapore

On-site
SGD 70,000 - 90,000