Site Reliability Engineer: Scale & Automate Global Services

TikTok USDS Joint Venture

Seattle (WA)

On-site

USD 130,000 - 246,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
401(k) with company match
Parental leave
Wellbeing benefits
Paid holidays

Job summary

TikTok USDS Joint Venture is seeking a Role in Seattle to help monitor and maintain availability for core TikTok services, ranging from video playback to recommendations. You will work with a global team to ensure reliability, scalability, and cost-efficiency of our distributed systems.

The role emphasizes hands-on management, incident response, and postmortems, with a focus on automating reliability improvements and maintaining high performance across the platform.

Qualifications

  • Bachelor or above degree in Computer Science or a related technical discipline with 3+ years experience in the deployment and administration of large-scale distributed systems
  • Strong understanding of Unix/Linux operating systems internals and administration, networking, storage systems, and database systems
  • Experience in programming languages such as C, C++, Java, Python, Go, Ruby, Rust, JavaScript
  • Experience in debugging and optimizing code and automate routine tasks
  • Experience with Nginx, Kubernetes, Docker, OpenStack, Hadoop, Spark, Flink, Kafka architectures
  • Experience in designing and analyzing large-scale distributed systems is preferred
  • Strong problem solving and communication skills

Responsibilities

  • Gaining a solid understanding of components and services powering TikTok experience
  • Maintain services to meet SLAs/SLOs by monitoring availability and performance
  • Support site-up issues as part of a global team for reliability and scalability
  • Scale systems sustainability via automation and efficiency improvements
  • Provide user support, incident responses and postmortems

Skills

C/C++
Java
Python
Go
Ruby
Rust
JavaScript
Distributed systems
Diagnostics & debugging
System design

Education

Bachelor's degree in Computer Science or related field

Tools

Nginx
Kubernetes
Docker
OpenStack
Hadoop
Spark
Flink
Kafka

Job description

TikTok USDS Joint Venture is seeking a Role in Seattle to help monitor and maintain availability for core TikTok services, ranging from video playback to recommendations. You will work with a global team to ensure reliability, scalability, and cost-efficiency of our distributed systems.

The role emphasizes hands-on management, incident response, and postmortems, with a focus on automating reliability improvements and maintaining high performance across the platform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — Automation, Scale & Uptime
Site Reliability Engineer — Automation, Scale & Uptime

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
Site Reliability Engineer — Video Platform at Scale
Site Reliability Engineer — Video Platform at Scale

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 113,000 - 178,000
Health insurance
401(k) match
Parental leave
+6
Site Reliability Engineer — Scalable Infra & Automation
Site Reliability Engineer — Scalable Infra & Automation

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
Senior Site Reliability Engineer - Scalable Cloud Systems
Senior Site Reliability Engineer - Scalable Cloud Systems

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 137,000 - 360,000
Site Reliability Engineer - Scale, Automation & Resilience
Site Reliability Engineer - Scale, Automation & Resilience

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 137,000 - 259,000
Video Platform SRE: Global Reliability Engineer
Video Platform SRE: Global Reliability Engineer

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
Medical insurance
401(k)
Paid parental leave
+2
Platform SRE Engineer — Scale, Availability & Incident Response
Platform SRE Engineer — Scale, Availability & Incident Response

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
Senior SRE Compute Platform: Scale & Reliability
Senior SRE Compute Platform: Scale & Reliability

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 178,000 - 342,000
SRE Tech Lead Manager: Observability & Reliability
SRE Tech Lead Manager: Observability & Reliability

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 198,000 - 416,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Site Reliability Engineer - Scale, Uptime & Observability
Site Reliability Engineer - Scale, Uptime & Observability

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 122,574 - 259,200
Medical, dental, and vision insurance
401(k) plan with company match
Paid parental leave
+6