Tech Lead (SRE) - Cloud Infrastructure

BYTEDANCE PTE. LTD.

Singapore

On-site

SGD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BYTEDANCE PTE. LTD. is seeking a Tech Lead for the Site Reliability Engineering (SRE) team.

You will guide software and systems engineers to design, build, and operate large-scale, distributed, and resilient infrastructure services. Emphasis on automation, reliability, and cross-team collaboration to support rapid improvement iterations. You will establish efficient project execution processes, promote sound engineering practices, and coordinate with other infrastructure teams and user

Qualifications

  • Bachelor's degree in Computer Science or closely related field and 5+ years of professional experience (including 3+ years in R&D)
  • Systematic approach to operations with proficiency in Linux and networking
  • Practical expertise in managing large-scale distributed systems
  • Self-motivated with strong planning and summarisation skills
  • Proven track record of project and team management
  • Experience with extensive cloud-computing platforms is a plus

Responsibilities

  • Establish and oversee the SRE team including recruitment, training, operations, coordination, and culture-building
  • Oversee development and acquisition of software systems; set long-term technical strategy with milestones
  • Develop PoC/solutions and provide technical leadership for software/platform features with security considerations
  • Create protocols for access management, disaster recovery, fault handling, and platform operations
  • Design and implement automated governance, monitoring, and SOA-based platforms
  • Collaborate with system development to ensure reliability from design to launch and improve automation
  • Foster cross-team communication and refine business processes; drive evolution of business architecture

Skills

Linux systems
Networking
Team leadership
Cloud platforms
Distributed systems

Education

Bachelor's degree in Computer Science or related field

Job description

Responsibilities

Team Introduction The Site Reliability Engineering (SRE) team is a fusion of software and systems engineering techniques used to design and operate large-scale, extensively distributed, and resilient systems. Within Infrastructure SRE, our primary focus is to ensure that the reliability and uptime of our infrastructure services meet the needs of our users and support rapid improvement iterations.


Our software development efforts are deeply committed to optimising existing systems, constructing essential infrastructure, and streamlining operations through automation. The Role In the role of a Tech Lead, you will assume responsibility for guiding and assembling a team of software and system engineers, leveraging your exceptional technical leadership skills.


Your role will involve establishing efficient processes for project execution and promoting sound engineering practices. Additionally, you will maintain regular coordination and communication with other infrastructure teams and our user community. What you will be doing:


1. Establish and oversee the SRE team, which encompasses tasks such as team recruitment, the training of new talent, system operation and maintenance, coordination efforts, and fostering a cohesive team culture; 2. Oversee the acquisition and development of software systems in organisational units. Establish a comprehensive long-term technical strategy with well-defined implementation steps and milestones to continually enhance the team's competitiveness and technological capabilities;


3. Oversee the development of Proof-of-Concept/solutions and provide technical expertise on the development of software and platform features, ensuring that appropriate security and risk factors are considered; 4. Create protocols and strategies for critical aspects of the operating platform, including access management, configuration, disaster recovery, and fault handling; 5. Devise and implement software platforms and monitoring frameworks that promote efficient, automated, and intelligent governance within a service-oriented architecture (SOA);


6. Collaborate closely with the system development team to guarantee the reliability of systems from initial design through to launch. Consistently advance automated operations and maintenance facilities and platforms; 7. Foster improved communication and collaboration with business teams, enhance cross-team coordination, and persistently refine and optimize business processes.


Drive the evolution of business architecture design.



Qualifications

Minimum Qualifications:



  • At least a Bachelor's Degree in Computer Science or a closely related technical field, along with more than 5 years of professional experience (including at least 3 years in Research and Development);

  • Demonstrates a systematic approach to operations and maintenance, with proficiency in Linux systems and networking.


Brings practical expertise in managing and maintaining large-scale distributed systems;



  • Self-motivated with strong planning and summarisation skills.


Possesses a track record of project and team management;



  • Exhibits a high level of responsibility, a proactive team-oriented attitude, and exceptional problem-solving abilities; Preferred Qualifications:

  • Prior experience with extensive cloud-computing platforms is a plus.

  • Experience in the development of large-scale distributed storage, scheduling, big data computing systems, or intelligent operations and maintenance.



Job Information

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.



Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect - and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.


As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an \"Always Day 1\" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.



Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend Software Engineer (SRE) - Cloud Infrastructure
Backend Software Engineer (SRE) - Cloud Infrastructure

ByteDance • Singapore

On-site
SGD 120,000 - 180,000
Cloud Site Relibility Engineer - DCS
Cloud Site Relibility Engineer - DCS

ByteDance • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer - Big Data Computer Platform
Site Reliability Engineer - Big Data Computer Platform

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Backend Software Engineer (SRE) Graduate (Cloud Infrastructure) - 2027 Start
Backend Software Engineer (SRE) Graduate (Cloud Infrastructure) - 2027 Start

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 67,000 - 112,000
Site Reliability Engineer, System - System Service Global
Site Reliability Engineer, System - System Service Global

ByteDance • Singapore

On-site
SGD 150,000 - 190,000
Site Reliability Engineer - Traffic Infrastructure
Site Reliability Engineer - Traffic Infrastructure

ByteDance • Singapore

On-site
SGD 150,000 - 210,000
Software Engineer (SRE - Platform Services) Intern (Infrastructure Engineering) - 2027 Start
Software Engineer (SRE - Platform Services) Intern (Infrastructure Engineering) - 2027 Start

ByteDance • Singapore

On-site
SGD 13,000 - 20,000
Site Reliability Engineer, System - System Service Global
Site Reliability Engineer, System - System Service Global

BYTEDANCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Backend Software Engineer (SRE) - Cloud Infrastructure Singapore Regular
Backend Software Engineer (SRE) - Cloud Infrastructure Singapore Regular

Bytedance • Singapore

On-site
SGD 110,000 - 180,000
Site Reliability Engineer Intern, System - System Service Global (Infrastructure Engineering), 2027 Start
Site Reliability Engineer Intern, System - System Service Global (Infrastructure Engineering), 2027 Start

ByteDance • Singapore

On-site
SGD 13,000 - 20,000