Site Reliability Engineer - Scale, Uptime & Observability
TikTok USDS Joint Venture
San Jose (CA)
On-site
USD 122,574 - 259,200
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Medical, dental, and vision insurance
401(k) plan with company match
Paid parental leave
Short-term and long-term disability coverage
Life insurance
Wellbeing benefits
10 paid holidays
10 paid sick days
17 days of paid personal time
Job summary
TikTok USDS Joint Venture is seeking a Site Reliability Engineer to design and optimize high-concurrency distributed systems, enhance observability, and lead incident response efforts. The suitable candidate should have a Bachelor's degree in Computer Science or related field and must be proficient in programming, with strong Linux system knowledge. Employees can expect a fully in-person work environment and generous benefits, including health insurance and paid time off.
Qualifications
Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience.
Proficiency in one or more programming languages (e.g., Go, Python, Java, or C++).
Strong understanding of Linux system internals and distributed systems.
Experience managing containerized environments such as Kubernetes or Docker.
Responsibilities
Design and optimize high-concurrency distributed systems.
Build and maintain automation tools to streamline deployments.
Develop monitoring, alerting, and logging systems.
Lead global disaster-recovery drills and validate failover mechanisms.
Respond to high-priority incidents and coordinate service restoration.
Facilitate blameless post-mortems and transform insights into requirements.
Manage capacity planning and resource allocation.
Skills
Proficiency in programming languages (Go, Python, Java, C++)
Bachelor’s degree in Computer Science or related field
Tools
Infrastructure as Code
Monitoring tools
Job description
TikTok USDS Joint Venture is seeking a Site Reliability Engineer to design and optimize high-concurrency distributed systems, enhance observability, and lead incident response efforts. The suitable candidate should have a Bachelor's degree in Computer Science or related field and must be proficient in programming, with strong Linux system knowledge. Employees can expect a fully in-person work environment and generous benefits, including health insurance and paid time off.