SRE for AI/ML Systems - Reliability & Scale Expert

TikTok USDS Joint Venture

Sydney

On-site

AUD 150,000 - 190,000

Full time

7 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

TikTok USDS Joint Venture LLC is seeking an on-site Site Reliability Engineer to design, build, and operate highly available distributed systems. You will monitor performance, automate incident response, and collaborate with software teams to ensure reliability and security across our AML AI/ML platforms.

The role requires strong Linux expertise, experience with distributed systems, and proficiency in languages like C/C++, Python, or Go.

Qualifications

  • Bachelor's or Master’s in Computer Science, Computer Engineering, or equivalent.
  • Experience analyzing and troubleshooting Linux-based distributed systems.
  • Knowledge of data structures and algorithms; relational databases.

Responsibilities

  • Design, build, and maintain highly available, scalable, and fault-tolerant systems.
  • Monitor and analyze system performance, identifying and resolving issues before user impact.
  • Develop and maintain automated monitoring, alerting, and incident response systems.
  • Collaborate with software engineering teams to ensure reliability, scalability, and performance in applications.
  • Implement security best practices and ensure regulatory compliance.
  • Participate in on-call rotations and respond to incidents during/off-hours.
  • Conduct root cause analysis and post-mortems to prevent recurrence.

Skills

Linux
Distributed systems
C/C++
Python
Go
Data structures
Relational databases

Education

Bachelor's/Master's in Computer Science/Computer Engineering or equivalent

Tools

TensorFlow
PyTorch
MXNet
PaddlePaddle

Job description

TikTok USDS Joint Venture LLC is seeking an on-site Site Reliability Engineer to design, build, and operate highly available distributed systems. You will monitor performance, automate incident response, and collaborate with software teams to ensure reliability and security across our AML AI/ML platforms.

The role requires strong Linux expertise, experience with distributed systems, and proficiency in languages like C/C++, Python, or Go.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer: Scale, Resilience & Automation
Senior Site Reliability Engineer: Scale, Resilience & Automation

TikTok USDS Joint Venture • Sydney

On-site
AUD 120,000 - 160,000
Senior SRE - Global Platform Reliability & Scale
Senior SRE - Global Platform Reliability & Scale

TikTok USDS Joint Venture • Sydney

On-site
AUD 150,000 - 190,000
Senior SRE Lead: AI Platform Reliability & Scale
Senior SRE Lead: AI Platform Reliability & Scale

TikTok • Sydney

On-site
AUD 180,000 - 250,000
Site Reliability Engineer: Scale, Automation & Infra Excellence
Site Reliability Engineer: Scale, Automation & Infra Excellence

TikTok • Sydney

On-site
AUD 150,000 - 210,000
Global E-commerce SRE: Scale, Uptime & Automation
Global E-commerce SRE: Scale, Uptime & Automation

TikTok USDS Joint Venture • Sydney

On-site
AUD 180,000 - 250,000
Security SRE Engineer: Platform Reliability & Automation
Security SRE Engineer: Platform Reliability & Automation

TikTok • Sydney

On-site
AUD 140,000 - 180,000
Site Reliability Engineer, TikTok Product
Site Reliability Engineer, TikTok Product

TikTok USDS Joint Venture • Sydney

On-site
AUD 120,000 - 160,000
Site Reliability Engineer, AML Platform - USDS
Site Reliability Engineer, AML Platform - USDS

TikTok USDS Joint Venture • Sydney

On-site
AUD 150,000 - 190,000
Senior Site Reliability Engineer, TikTok Product
Senior Site Reliability Engineer, TikTok Product

TikTok USDS Joint Venture • Sydney

On-site
AUD 150,000 - 190,000
Senior Site Reliability Engineer - Infrastructure
Senior Site Reliability Engineer - Infrastructure

TikTok • Sydney

On-site
AUD 180,000 - 250,000