Site Reliability Engineer - Traffic Infrastructure

Bytedance

Singapore

On-site

SGD 120,000 - 180,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

ByteDance seeks a seasoned SRE to own operation, stability, and delivery of the Network-Traffic Infrastructure. You will design observability, incident response, and emergency systems for a global edge platform, working across China and non-China regions.

The role requires hands-on experience with Kubernetes, cloud networking and microservices, plus strong analytical and communication skills to collaborate with distributed teams. Singapore-based position with growth in a fast-paced tech company.

Qualifications

  • Bachelor's degree or above in computer science or a related field, with at least 3 years of relevant experience in R&D, system operation and maintenance, or SRE.
  • Solid understanding of Kubernetes, edge computing, cloud networking, Load Balancing, microservice architecture and related technologies.
  • Strong analytical skills, excellent communication abilities, a strong sense of responsibility and team spirit.

Responsibilities

  • Operation and maintenance and stability assurance of ByteDance's Network-Traffic Infrastructure.
  • Delivery and operation & maintenance of production system, including delivery, change and release of facilities, components and products.
  • Design and implementation of stability assurance system (monitoring/alerts/logging, RCA/impact assessment, issue resolution).
  • Design and implementation of emergency response system, including ticket processing, emergency response, risk governance, and long-term optimization.

Skills

Kubernetes
Edge computing
Cloud networking
Load balancing
Microservices
Analytical skills
Teamwork

Education

Bachelor's degree in CS or related field

Tools

Prometheus
Grafana
Cloud platforms (AWS/GCP/Azure)

Job description

About the Team

The Traffic Infrastructure team leverages unified platform capabilities to manage global edge infrastructure (China & Non-China), both self-built and third-party, providing standardized, compliant, scalable, and cost-effective traffic infrastructure capabilities for edge services. Our vision is to build a global edge traffic infrastructure platform and become the long-term cornerstone of ByteDance’s global edge business in terms of scale, performance, and cost.

Responsibilities
  • Responsible for the operation and maintenance as well as stability assurance of ByteDance's "Network-Traffic Infrastructure".
  • Responsible for the delivery and operation & maintenance of the production system, including the delivery, change and release of facilities, components and products, and improving the efficiency of both delivery and operation & maintenance.
  • Responsible for the design and implementation of the stability assurance system, covering system observability (monitoring/alerts/logging), troubleshooting (root cause analysis/impact assessment), and issue resolution (manual/self-healing).
  • Responsible for the design and implementation of the emergency response system, including work such as ticket processing, emergency response, risk governance, and long-term optimization, to enhance the risk emergency response capability.
Qualifications
Minimum Qualification(s)
  • Bachelor's degree or above in computer science or a related field, with at least 3 years of relevant experience in R&D, system operation and maintenance, or SRE.
  • Familiar with infrastructure architecture, and have a solid understanding of Kubernetes, edge computing, cloud networking, Load Balance, microservice architecture and other related technologies.
  • Possess strong analytical skills, excellent communication abilities, a strong sense of responsibility and team spirit.
Preferred Qualification(s)
  • Possess practical operation and maintenance and stability assurance experience in Kubernetes, cloud computing, edge computing, and cloud networking.
  • Well-versed in high availability, stability assurance, and emergency response systems for infrastructure or distributed systems, with relevant operation and maintenance experience.
  • Hands-on experience in handling risks, potential hazards and failures of infrastructure or distributed systems.
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives.

Our platform connects people from across the globe and so does our workplace.

At ByteDance, our mission is to inspire creativity and enrich life.

To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach.

We are passionate about this and hope you are too.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Operations and Platform Engineer - Traffic Infrastructure
Operations and Platform Engineer - Traffic Infrastructure

Bytedance • Singapore

On-site
SGD 120,000 - 190,000
Site Reliability Engineer - Traffic Infrastructure Technology - Infrastructure Singapore Regular
Site Reliability Engineer - Traffic Infrastructure Technology - Infrastructure Singapore Regular

Bytedance • Singapore

On-site
SGD 120,000 - 180,000
Backend Software Engineer (Cloud Platform) - Traffic Infrastructure
Backend Software Engineer (Cloud Platform) - Traffic Infrastructure

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer - Big Data Computer Platform
Site Reliability Engineer - Big Data Computer Platform

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer, System - System Service Global
Site Reliability Engineer, System - System Service Global

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Backend Software Engineer (Cloud Platform) - Traffic Infrastructure
Backend Software Engineer (Cloud Platform) - Traffic Infrastructure

ByteDance • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer, Traffic Solution - System Service Global
Site Reliability Engineer, Traffic Solution - System Service Global

ByteDance • Singapore

On-site
SGD 90,000 - 140,000
Site Reliability Engineer, System - System Service Global
Site Reliability Engineer, System - System Service Global

ByteDance • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer (Cloud) - Infrastructure Engineering Technology - DevOps Singapore Regular
Site Reliability Engineer (Cloud) - Infrastructure Engineering Technology - DevOps Singapore Regular

ByteDance • Singapore

On-site
SGD 60,000 - 90,000
Cloud Site Relibility Engineer - DCS
Cloud Site Relibility Engineer - DCS

ByteDance • Singapore

On-site
SGD 110,000 - 180,000