Tech Lead Cloud Site Reliability Engineer - DCS Cloud

ByteDance

San Jose (CA)

On-site

USD 244,800 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance in San Jose seeks a Tech Lead Cloud Site Reliability Engineer to design, build, and operate a global infrastructure spanning public and private clouds. You’ll drive automation, create AMIs, and guide incident response in a fast-paced environment.

You will lead improvements across the full lifecycle, contribute to Cloud Host Delivery, Operation, and Security groups, and partner with cross‑functional teams. A strong background in Linux, networking, GPUs, and cloud-native tech is required.

Qualifications

  • Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • 5+ years of experience in Linux operations, SRE, or DevOps.
  • Proficiency in Go, Python, or C++, with strong automation skills.
  • Deep understanding of Linux OS, networks, storage, GPUs, databases, and troubleshooting.
  • Familiarity with monitoring, alerting, capacity, change, canary releases, incidents, postmortems.
  • Excellent communication and collaboration skills, ownership mindset.
  • Preferred: experience with public clouds (OCI, AWS, Azure, GCP).
  • Experience with image/AMI systems, resource scheduling, virtualization (KVM/QEMU).
  • Knowledge of containers and cloud-native ecosystems (Docker, Kubernetes, containerd).
  • GPU clusters: drivers, CUDA, MIG, topology, stress testing, delivery pipelines.
  • Proven reliability initiatives: failure drills, capacity governance, observability, cost optimization.
  • Open-source contributions or patents highly regarded.
  • Experience operating large-scale production environments.

Responsibilities

  • Design, build, scale, and operate ByteDance’s global infrastructure spanning public and private clouds.
  • Develop tools, automation frameworks, visualizations, and monitoring to optimize operations.
  • Create and manage cloud AMIs/images across environments with strict compliance.
  • Engage in fast-paced on-call rotations addressing incidents in cloud, OS, network, and performance.
  • Lead improvements across the full infrastructure lifecycle from ideation to deployment and support.
  • Contribute to subgroups: Cloud Host Delivery – Delivery & Standardization; Cloud Host Operation – Efficiency & Reliability; Cloud Management & Security.

Skills

Go
Python
C++
Linux basics
Networking basics
Troubleshooting
Communication
On-call/Incidents

Education

Bachelor's degree in CS/related field

Tools

Docker
Kubernetes
containerd
KVM/QEMU
AWS
OCI
Git

Job description

Tech Lead Cloud Site Reliability Engineer - DCS Cloud

Location: San Jose

Team: Technology

Employment Type: Regular

Job Code: A14024C

Responsibilities
  • Design, build, scale, and operate ByteDance’s global infrastructure, including large‑scale systems spanning public and private clouds.
  • Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure.
  • Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the company’s global compliance standards.
  • Engage in a fast‑paced environment, participating in technical operations and on‑call rotations to address incidents related to cloud, OS, network, performance, and reliability.
  • Lead improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.
  • Contribute to the following subgroups:
    • Cloud Host Delivery – Delivery & Standardization
    • Cloud Host Operation – Operation Efficiency & Reliability
    • Cloud Management & Security
Qualifications
  • Minimum: Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • 5+ years of experience in Linux operations, SRE, or DevOps.
  • Proficiency in at least one programming language such as Go, Python, or C++ with strong platform‑development and automation skills.
  • Deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, databases, and systematic troubleshooting/root‑cause analysis.
  • Familiarity with core reliability practices: monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and postmortem processes.
  • Excellent communication and collaboration skills, proactive problem identification, cross‑team execution, and ownership mindset.
  • Preferred: Experience operating public cloud platforms (OCI, AWS, Azure, GCP, etc.) and large‑scale cloud host delivery.
  • Experience with image/AMI systems, resource scheduling, network adaptation, virtualization technologies such as KVM/QEMU.
  • Knowledge of containers and cloud‑native ecosystems (Docker, Kubernetes, containerd) including isolation mechanisms (cgroups, namespaces).
  • Experience maintaining GPU clusters: drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, GPU delivery pipelines.
  • Proven reliability‑focused initiatives: failure drill systems, capacity governance, change governance, observability platforms, resource cost optimization.
  • Open‑source contributions, technical blogs, patents, or technical sharing preferred.
  • Experience operating large‑scale production environments.
Job Information

The base salary range for this position in the selected city is $244,800 - $450,000 annually. Compensation may vary outside of this range depending on factors such as qualifications, skills, competencies, and experience. Base pay is one part of the Total Package, and this role may be eligible for additional discretionary bonuses/incentives and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance; a 401(k) savings plan with company match; paid parental leave; short‑term and long‑term disability coverage; life insurance; wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year, and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The company reserves the right to modify or change these benefits programs at any time, with or without notice.

Equal Opportunity Statement

For Los Angeles County (unincorporated) candidates: Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse, and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

  • Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  • Appropriately handling and managing confidential information, including proprietary and trade secret information and access to information technology systems;
  • Exercising sound judgment.

Reasonable Accommodation: ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs, or other reasons protected by applicable laws. If you require assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Site Reliability Engineer - DCS Cloud San Jose Regular
Cloud Site Reliability Engineer - DCS Cloud San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 136,000 - 360,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid personal time off
+1
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
Tech Lead Cloud Site Reliability Engineer - DCS Cloud

Socket.dev • Seattle (WA)

On-site
USD 232,560 - 427,500
Cloud Site Reliability Engineer - DCS Cloud
Cloud Site Reliability Engineer - DCS Cloud

ByteDance • San Jose (CA)

On-site
USD 136,000 - 360,000
Medical, dental, and vision insurance
401(k) plan with company match
Paid parental leave
+4
Cloud Site Reliability Engineer - DCS Cloud Seattle Regular
Cloud Site Reliability Engineer - DCS Cloud Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 129,000 - 342,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+5
Tech Lead,Infrastructure Delivery Platform
Tech Lead,Infrastructure Delivery Platform

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Backend Software Engineer - Platforms
Backend Software Engineer - Platforms

ByteDance • New York (NY)

On-site
USD 156,000 - 388,000
Tech Lead, Research Scientist - DPU & AI Infra
Tech Lead, Research Scientist - DPU & AI Infra

ByteDance • San Jose (CA)

On-site
USD 244,800 - 588,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
Senior Production System Engineer - San Jose
Senior Production System Engineer - San Jose

ByteDance • San Jose (CA)

On-site
USD 115,000 - 288,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
System Engineer (Operating System) - System Technologies and Engineering
System Engineer (Operating System) - System Technologies and Engineering

ByteDance • San Jose (CA)

On-site
USD 156,000 - 388,000
Medical insurance
Dental insurance
Vision insurance
+9
Cloud Site Reliability Engineer - DCS Cloud
Cloud Site Reliability Engineer - DCS Cloud

ByteDance • Seattle (WA)

On-site
USD 148,000 - 301,000