Site Reliability Engineer, Reliability Team - USDS

TikTok USDS Joint Venture

San Jose (CA)

On-site

USD 122,574 - 259,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) plan with company match
Paid parental leave
Short-term and long-term disability coverage
Life insurance
Wellbeing benefits
10 paid holidays
10 paid sick days
17 days of paid personal time

Job summary

TikTok USDS Joint Venture is seeking a Site Reliability Engineer to design and optimize high-concurrency distributed systems, enhance observability, and lead incident response efforts. The suitable candidate should have a Bachelor's degree in Computer Science or related field and must be proficient in programming, with strong Linux system knowledge. Employees can expect a fully in-person work environment and generous benefits, including health insurance and paid time off.

Qualifications

  • Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience.
  • Proficiency in one or more programming languages (e.g., Go, Python, Java, or C++).
  • Strong understanding of Linux system internals and distributed systems.
  • Experience managing containerized environments such as Kubernetes or Docker.

Responsibilities

  • Design and optimize high-concurrency distributed systems.
  • Build and maintain automation tools to streamline deployments.
  • Develop monitoring, alerting, and logging systems.
  • Lead global disaster-recovery drills and validate failover mechanisms.
  • Respond to high-priority incidents and coordinate service restoration.
  • Facilitate blameless post-mortems and transform insights into requirements.
  • Manage capacity planning and resource allocation.

Skills

Proficiency in programming languages (Go, Python, Java, C++)
Understanding of Linux system internals
Experience managing containerized environments (Kubernetes, Docker)

Education

Bachelor’s degree in Computer Science or related field

Tools

Infrastructure as Code
Monitoring tools

Job description

Role Overview

The Site Reliability Engineering (SRE) team at TikTok builds and operates large‑scale, fault‑tolerant systems that power TikTok’s core services worldwide. The team focuses on enhancing observability and operability, using data insights to maintain 24/7 business stability.

Responsibilities
  • Design and optimize high‑concurrency distributed systems, collaborating with development teams to ensure scalability, reliability, and high availability.
  • Build and maintain automation tools to reduce toil, streamline deployments, and manage infrastructure as code.
  • Develop and refine monitoring, alerting, and logging systems (SLIs/SLOs) for deep visibility into service health and performance.
  • Lead global disaster‑recovery drills, simulate complex failure scenarios, and validate failover mechanisms to keep the platform operational under extreme conditions.
  • Respond to high‑priority incidents as a key responder, coordinating cross‑functional war rooms, driving technical troubleshooting, and leading service restoration.
  • Facilitate blameless post‑mortems and root‑cause analysis; transform incident insights into engineering requirements to harden systems.
  • Manage capacity planning, resource allocation, and performance bottlenecks to accommodate organic growth and traffic surges.
Minimum Qualifications
  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • Proficiency in one or more programming languages (e.g., Go, Python, Java, or C++).
  • Strong understanding of Linux system internals, networking (TCP/IP, DNS, load balancing), and distributed systems.
  • Experience managing containerized environments such as Kubernetes or Docker.
Preferred Qualifications
  • Experience in a high‑traffic production environment with a focus on incident response and site stability.
  • Hands‑on knowledge of disaster‑recovery strategies, including multi‑region failover and data consistency in distributed databases.
  • Familiarity with observability and monitoring tools.
  • Experience with Infrastructure as Code.
Benefits & Compensation

Compensation for this position ranges from $122,574 to $259,200 annually, with potential bonuses, incentives, and restricted stock units. Benefits include medical, dental, and vision insurance; a 401(k) plan with company match; paid parental leave; short‑term and long‑term disability coverage; life insurance; and wellbeing benefits. Employees receive 10 paid holidays, 10 paid sick days, and 17 days of paid personal time (prorated on hire and increasing with tenure).

Work Environment

The company is shifting to a fully in‑person schedule, up to five days a week, to enhance speed, alignment, and agility across teams.

Reasonable Accommodation

USDS is committed to providing reasonable accommodations for candidates with disabilities, pregnancy, sincerely held religious beliefs, or other protected reasons. Assistance can be requested at https://tinyurl.com/USDS-RA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Infrastructure and Assurance Services - USDS
Site Reliability Engineer, Infrastructure and Assurance Services - USDS

TikTok • Seattle (WA)

Hybrid
USD 112,000 - 178,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+6
Site Reliability Engineer, Infrastructure and Assurance Services - USDS
Site Reliability Engineer, Infrastructure and Assurance Services - USDS

TikTok • Seattle (WA)

On-site
USD 100,000 - 150,000
Site Reliability Engineer, Product - USDS
Site Reliability Engineer, Product - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 122,000 - 260,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
Site Reliability Engineer, Compute - USDS
Site Reliability Engineer, Compute - USDS

TikTok • Seattle (WA)

On-site
USD 100,000 - 130,000
Medical, dental, and vision insurance
401(k) with company match
Parental leave
+1
Senior Site Reliability Engineer, Network and Traffic - USDS
Senior Site Reliability Engineer, Network and Traffic - USDS

TikTok • San Jose (CA)

On-site
USD 136,000 - 360,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Site Reliability Engineer - Video Platform - USDS
Site Reliability Engineer - Video Platform - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
Site Reliability Engineer - AML Global Recommendation - USDS
Site Reliability Engineer - AML Global Recommendation - USDS

TikTok • San Jose (CA)

Hybrid
USD 118,000 - 260,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Site Reliability Engineer - Video Platform - (Entry Level)- USDS
Site Reliability Engineer - Video Platform - (Entry Level)- USDS

TikTok • San Jose (CA)

On-site
USD 118,000 - 188,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+4
Site Reliability Engineer, Cloud Infrastructure - USDS
Site Reliability Engineer, Cloud Infrastructure - USDS

TikTok • San Jose (CA)

Hybrid
USD 118,000 - 260,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
SRE Tech Lead Manager, Product - USDS
SRE Tech Lead Manager, Product - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 208,000 - 438,000