Site Reliability Engineer - AI Application

ByteDance

Singapore

On-site

SGD 120,000 - 170,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a senior Reliability/Platform engineer to ensure the continuous operation of core Viking team systems. You will focus on capacity planning, system monitoring, and rapid fault localization for AI search and vector databases.

The role requires strong Linux fundamentals, distributed systems knowledge, and experience with Kubernetes/Docker/OpenStack. Collaboration across global teams is essential for high-availability architecture improvements.

Qualifications

  • Bachelor's degree or above in computer-related fields; 5+ years of relevant work experience.
  • Solid foundation in Linux OS, storage, and network IO principles.
  • Familiar with at least one programming language (Python/Go/Java/Shell/Ansible) with decent operations mindset.
  • Knowledge of cloud infrastructure (AWS/Volcano Engine/Aliyun/GCP) and distributed systems (Nginx/Kubernetes/Docker/OpenStack/Hadoop/Spark/Flink).

Responsibilities

  • Ensure reliability and operation of core Viking Team big data and online services; plan capacity and stability.
  • Improve system visibility by monitoring availability and performance metrics; support quick fault location, ensure AI search/vector DB links run smoothly.
  • Enhance reliability, scalability, and performance to meet core SLA targets.
  • Contribute to design and implementation of automation platforms for large-scale Viking clusters and AI search-related clusters.
  • Analyze performance bottlenecks and governance; aid high-availability architecture upgrades.

Skills

Linux proficiency
Programming languages
Cloud infrastructure
Distributed systems

Education

Bachelor's degree or above in computer-related fields

Tools

Nginx
Kubernetes
Docker
OpenStack
Hadoop
Spark
Flink
AWS
GCP

Job description

Responsibilities

About the team

We are an AI-driven search and recommendation team focused on building innovative, scalable products for global users.

  1. 1. Ensure the reliability and normal operation of multiple core systems related to Viking Team's Big data and online services, while focusing on system capacity planning and stability assurance;
  2. 2. Enhance system visibility by monitoring the availability and performance metrics of system components, helping development teams quickly locate faults, and especially ensuring operation of critical links such as AI search/vector databases;
  3. 3. Improve the reliability, scalability, and Performance optimization of services to ensure the achievement of the core system SLA;
  4. 4. Participated in the design and implementation of the automation platform, ensuring the rapid iteration and efficient operation and maintenance of large-scale online Viking clusters and AI search-related clusters;
  5. 5. Combining with the usage scenarios of AI Search/Viking business, in-depth optimization of service governance practices, including but not limited to analysis of performance bottlenecks in key AI Search/Viking links, business problem location and troubleshooting, promoting the transformation and upgrading of the system's high-availability architecture, and those familiar with Viking-related technologies are preferred to participate in core optimization work.
Qualifications
Minimum Qualifications
  1. 1. Bachelor's degree or above, majoring in computer-related fields, with more than five years of relevant work experience;
  2. 2. Has a solid foundation in computer software knowledge, and understands the relevant principles of Linux operating systems, storage, network IO, etc.
  3. 3. Familiar with at least one programming language (such as Python/Go/Java/Shell/Ansible), with moderate development capabilities, and placing more emphasis on operations and maintenance practices and problem-solving abilities;
  4. 4. Understand at least one type of knowledge related to cloud infrastructure such as AWS/Volcano Engine/Aliyun/GCP; those with experience in computing/distributed systems are preferred (e.g., Nginx/Kubernetes/Docker/OpenStack/Hadoop/Spark/Flink, etc.);
Preferred Qualifications
  1. 1. Familiar with algorithmic thinking, good data structure and system design capabilities
  2. 2. Have certain understanding of AI Cloud, large model-related Search Suggestion, and Recommender system.
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - AI Application Technology - Backend Singapore Regular
Site Reliability Engineer - AI Application Technology - Backend Singapore Regular

ByteDance • Singapore

On-site
SGD 80,000 - 120,000
Innovative work environment
Diversity and inclusion initiatives
Opportunity for career growth
Cloud Site Relibility Engineer - DCS
Cloud Site Relibility Engineer - DCS

ByteDance • Singapore

On-site
SGD 90,000 - 150,000
Backend Software Engineer (SRE) - Cloud Infrastructure
Backend Software Engineer (SRE) - Cloud Infrastructure

ByteDance • Singapore

On-site
SGD 120,000 - 180,000
Machine Learning System Engineer- Data AML- Soaring Star Talent Program
Machine Learning System Engineer- Data AML- Soaring Star Talent Program

Pangleglobal • Singapore

On-site
SGD 80,000 - 120,000
Software Engineer - Service Platform
Software Engineer - Service Platform

ByteDance • Singapore

On-site
SGD 120,000 - 180,000
Senior Backend Software Engineer (AI Infrastructure / Artifact Management) Developer Services
Senior Backend Software Engineer (AI Infrastructure / Artifact Management) Developer Services

Bytedance • Singapore

On-site
SGD 120,000 - 160,000
Applied AI Architect, BytePlus
Applied AI Architect, BytePlus

ByteDance • Singapore

On-site
SGD 120,000 - 180,000
Graduate ML System Engineer - Large-Scale Systems
Graduate ML System Engineer - Large-Scale Systems

ByteDance • Singapore

On-site
SGD 60,000 - 80,000
Senior Systems Engineer – Server Provisioning & Deployment, DCS
Senior Systems Engineer – Server Provisioning & Deployment, DCS

ByteDance • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer Intern, System - System Service Global (Infrastructure Engineering), 2027 Start
Site Reliability Engineer Intern, System - System Service Global (Infrastructure Engineering), 2027 Start

ByteDance • Singapore

On-site
SGD 13,000 - 20,000