SRE for Large-Scale AI/ML Systems

TikTok USDS Joint Venture

Seattle (WA)

On-site

USD 130,000 - 246,000

Full time

47 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Parental leave
Wellbeing benefits
Paid holidays

Job summary

TikTok USDS Joint Venture seeks a Site Reliability Engineer for the Applied Machine Learning team to design, build, and operate massively distributed AI/ML systems such as recommender models and large language models in the United States and globally.

You will sharpen skills in coding, performance analysis, and large-scale system operation while influencing hardware/capacity decisions and automating monitoring, alerts, and incident response across global platforms.

Qualifications

  • Expertise in analyzing and troubleshooting Linux-based distributed systems.
  • Bachelor's/Master's degree in Computer Science, Computer Engineering, or equivalent years of experience in a SRE or software engineering role.
  • Experience programming with C, C++, Python, or Go.
  • Strong understanding of data structures and algorithms.
  • Competent knowledge of relational database systems.

Responsibilities

  • Design, build, and maintain highly available, scalable, and fault-tolerant systems.
  • Monitor and analyze system performance, identifying and resolving issues before user impact.
  • Develop and maintain automated monitoring, alerting, and incident response systems.
  • Collaborate with software engineering teams to ensure reliability, scalability, and performance.
  • Implement and maintain security best practices and regulatory compliance.
  • Participate in on-call rotations and respond to incidents during/after business hours.
  • Conduct root cause analysis of incidents and implement preventative measures.

Skills

Linux
C/C++
Python
Go
Relational DB
Algorithms

Education

Bachelor's/Master's in CS/CE or equivalent

Tools

TensorFlow
PyTorch
MXNet
PaddlePaddle

Job description

TikTok USDS Joint Venture seeks a Site Reliability Engineer for the Applied Machine Learning team to design, build, and operate massively distributed AI/ML systems such as recommender models and large language models in the United States and globally.

You will sharpen skills in coding, performance analysis, and large-scale system operation while influencing hardware/capacity decisions and automating monitoring, alerts, and incident response across global platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE, Applied ML — Scale, Reliability & Automation
SRE, Applied ML — Scale, Reliability & Automation

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
SRE Engineer, AI Infrastructure — Scale & Automation
SRE Engineer, AI Infrastructure — Scale & Automation

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 129,000 - 247,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Tech Lead AI Infra SRE — Scale Global Reliability
Tech Lead AI Infra SRE — Scale Global Reliability

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 198,000 - 416,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
Senior SRE Compute Platform: Scale & Reliability
Senior SRE Compute Platform: Scale & Reliability

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 178,000 - 342,000
SRE: AI-Driven Platform Reliability & Automation
SRE: AI-Driven Platform Reliability & Automation

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
Medical Insurance
Dental Insurance
Vision Insurance
+8
Tech Lead, Large-Scale ML Infra & Distributed Systems
Tech Lead, Large-Scale ML Infra & Distributed Systems

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 198,000 - 416,000
Site Reliability Engineer — Scalable Infra & Automation
Site Reliability Engineer — Scalable Infra & Automation

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
AI Infra SRE Engineer: Scale & Automation
AI Infra SRE Engineer: Scale & Automation

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer - Scalable Cloud Systems
Senior Site Reliability Engineer - Scalable Cloud Systems

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 137,000 - 360,000
Site Reliability Engineer - Scale, Automation & Resilience
Site Reliability Engineer - Scale, Automation & Resilience

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 137,000 - 259,000