Cloud SRE Tech Lead - Global Infra & Automation

ByteDance

San Jose (CA)

On-site

USD 244,800 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance in San Jose seeks a Tech Lead Cloud Site Reliability Engineer to design, build, and operate a global infrastructure spanning public and private clouds. You’ll drive automation, create AMIs, and guide incident response in a fast-paced environment.

You will lead improvements across the full lifecycle, contribute to Cloud Host Delivery, Operation, and Security groups, and partner with cross‑functional teams. A strong background in Linux, networking, GPUs, and cloud-native tech is required.

Qualifications

  • Bachelor’s degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • 5+ years of experience in Linux operations, SRE, or DevOps.
  • Proficiency in Go, Python, or C++, with strong automation skills.
  • Deep understanding of Linux OS, networks, storage, GPUs, databases, and troubleshooting.
  • Familiarity with monitoring, alerting, capacity, change, canary releases, incidents, postmortems.
  • Excellent communication and collaboration skills, ownership mindset.
  • Preferred: experience with public clouds (OCI, AWS, Azure, GCP).
  • Experience with image/AMI systems, resource scheduling, virtualization (KVM/QEMU).
  • Knowledge of containers and cloud-native ecosystems (Docker, Kubernetes, containerd).
  • GPU clusters: drivers, CUDA, MIG, topology, stress testing, delivery pipelines.
  • Proven reliability initiatives: failure drills, capacity governance, observability, cost optimization.
  • Open-source contributions or patents highly regarded.
  • Experience operating large-scale production environments.

Responsibilities

  • Design, build, scale, and operate ByteDance’s global infrastructure spanning public and private clouds.
  • Develop tools, automation frameworks, visualizations, and monitoring to optimize operations.
  • Create and manage cloud AMIs/images across environments with strict compliance.
  • Engage in fast-paced on-call rotations addressing incidents in cloud, OS, network, and performance.
  • Lead improvements across the full infrastructure lifecycle from ideation to deployment and support.
  • Contribute to subgroups: Cloud Host Delivery – Delivery & Standardization; Cloud Host Operation – Efficiency & Reliability; Cloud Management & Security.

Skills

Go
Python
C++
Linux basics
Networking basics
Troubleshooting
Communication
On-call/Incidents

Education

Bachelor's degree in CS/related field

Tools

Docker
Kubernetes
containerd
KVM/QEMU
AWS
OCI
Git

Job description

ByteDance in San Jose seeks a Tech Lead Cloud Site Reliability Engineer to design, build, and operate a global infrastructure spanning public and private clouds. You’ll drive automation, create AMIs, and guide incident response in a fast-paced environment.

You will lead improvements across the full lifecycle, contribute to Cloud Host Delivery, Operation, and Security groups, and partner with cross‑functional teams. A strong background in Linux, networking, GPUs, and cloud-native tech is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead, Data Infrastructure & SRE (Cloud-Scale)
Tech Lead, Data Infrastructure & SRE (Cloud-Scale)

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Health insurance
401(k) with company match
Paid parental leave
+3
Tech Lead - DevInfra Platform (Cloud IDE & CI/CD)
Tech Lead - DevInfra Platform (Cloud IDE & CI/CD)

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical Insurance
401k Matching
Parental Leave
+6
Senior SRE - Data Infrastructure & Reliability
Senior SRE - Data Infrastructure & Reliability

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
Tech Lead Cloud Site Reliability Engineer - DCS Cloud

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Senior Cloud Infrastructure Engineer - Scale & Automate
Senior Cloud Infrastructure Engineer - Scale & Automate

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical, dental and vision insurance
401(k) with company match
Parental leave
+6
Global Data Center SRE | Reliable Infrastructure
Global Data Center SRE | Reliable Infrastructure

ByteDance • San Jose (CA)

On-site
USD 210,000 - 330,000
Infra Delivery Platform Tech Lead — Architect & Team Leader
Infra Delivery Platform Tech Lead — Architect & Team Leader

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Graduate Backend Engineer — Global Infrastructure & Ops
Graduate Backend Engineer — Global Infrastructure & Ops

Bytedance • San Jose (CA)

On-site
USD 110,000 - 160,000
Tech Lead: Data Infrastructure & Site Reliability
Tech Lead: Data Infrastructure & Site Reliability

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500
Medical insurance
Dental and vision insurance
401(k) with company match
+7
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
Tech Lead Cloud Site Reliability Engineer - DCS Cloud

Socket.dev • Seattle (WA)

On-site
USD 232,560 - 427,500