SRE: AI-Driven Platform Reliability & Automation

TikTok USDS Joint Venture

Seattle (WA)

On-site

USD 130,000 - 246,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
401(k) Plan
Parental Leave
Disability Insurance
Life Insurance
Wellbeing Benefits
Paid Holidays
Paid Sick Days
Paid Personal Time

Job summary

TikTok USDS Joint Venture LLC is seeking an engineer to manage data services and real-time/batch pipelines, ensuring reliability and performance at scale. You will design AI-enabled automation to streamline incident response and monitoring, while enhancing operational efficiency.

The role requires Unix/Linux expertise, distributed systems knowledge, and hands-on experience with monitoring tools. This position includes 24/7 on-call duties with rotating shifts to support global users.

Qualifications

  • Bachelor or above degree in computer science or a related technical discipline.
  • At least 1 year of industrial experience.
  • Experience integrating AI/LLM APIs into internal workflows or infrastructure tooling.
  • Demonstrated independent thinking capabilities and troubleshooting skills.
  • Familiar with Unix/Linux system internals, networking, and distributed systems.
  • Expertise in monitoring tools (Prometheus, Grafana, DataDog) and fundamental observability approaches.

Responsibilities

  • Manage day-to-day operations of data service and real-time/batch data pipelines, including SLA/SLO/SLI management, deployment, performance tuning, and troubleshooting.
  • Design and deploy AI agents and LLM-powered automation to streamline incident response and proactive monitoring.
  • Create tools to improve system administration and operational efficiency using AI-assisted development tools.
  • Oversee the full lifecycle of services from inception, design, development, capacity planning, to deployment and refinement.
  • Support incident response and post-mortems with a focus on sustainability and reliability.
  • Work as part of a 24/7 team with rotating shifts, including holidays.

Skills

Data pipelines
AI/LLM integration
Monitoring & observability
Unix/Linux knowledge
Distributed systems

Education

Bachelor or above in CS/related field

Tools

MySQL
Redis
Nginx
Kafka
Kubernetes
Docker
Hadoop/Spark/Flink/Hive/OLAP/ClickHouse

Job description

TikTok USDS Joint Venture LLC is seeking an engineer to manage data services and real-time/batch pipelines, ensuring reliability and performance at scale. You will design AI-enabled automation to streamline incident response and monitoring, while enhancing operational efficiency.

The role requires Unix/Linux expertise, distributed systems knowledge, and hands-on experience with monitoring tools. This position includes 24/7 on-call duties with rotating shifts to support global users.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE, Applied ML — Scale, Reliability & Automation
SRE, Applied ML — Scale, Reliability & Automation

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
SRE Engineer, AI Infrastructure — Scale & Automation
SRE Engineer, AI Infrastructure — Scale & Automation

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 129,000 - 247,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Senior SRE Compute Platform: Scale & Reliability
Senior SRE Compute Platform: Scale & Reliability

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 178,000 - 342,000
SRE for Large-Scale AI/ML Systems
SRE for Large-Scale AI/ML Systems

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
Health insurance
401(k) matching
Parental leave
+2
Platform SRE Engineer — Scale, Availability & Incident Response
Platform SRE Engineer — Scale, Availability & Incident Response

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
AI Infra SRE Engineer: Scale & Automation
AI Infra SRE Engineer: Scale & Automation

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 120,000 - 180,000
Tech Lead AI Infra SRE — Scale Global Reliability
Tech Lead AI Infra SRE — Scale Global Reliability

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 198,000 - 416,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
Data Infrastructure SRE: Reliability & AI Ops
Data Infrastructure SRE: Reliability & AI Ops

TikTok • San Jose (CA)

On-site
USD 162,000 - 317,000
Video Platform SRE: Global Reliability Engineer
Video Platform SRE: Global Reliability Engineer

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
Medical insurance
401(k)
Paid parental leave
+2
Senior Data Infra SRE: Reliability, Scale & Automation
Senior Data Infra SRE: Reliability, Scale & Automation

TikTok • Seattle (WA)

On-site
USD 207,000 - 368,000