Site Reliability Engineer: Scale, Automation

TikTok USDS Joint Venture

Seattle (WA)

On-site

USD 130,000 - 246,000

Full time

34 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TikTok USDS Joint Venture invites a Site Reliability Engineer to join a team building and operating large-scale, distributed, fault-tolerant systems. You’ll design automation, monitor health, and contribute to on-call incident management across cross-functional teams.

Based in Seattle area, you’ll apply Go and Python skills, work with Linux, cloud platforms, and container tech like Docker and Kubernetes to scale services while upholding reliability and performance standards.

Qualifications

  • Bachelor's degree in CS, IT, or related field with 3+ years of experience.
  • Proven SRE/Systems Engineer experience.
  • Automation focus with Go or Python.
  • Experience with Linux and open-source technologies.
  • Experience with cloud and large-scale distributed systems.
  • Excellent problem-solving and cross-team collaboration.

Responsibilities

  • Develop and maintain automation procedures to maximize system efficiency and minimize human intervention.
  • Work closely with software engineering teams to design, deploy and operate elements to ensure systems are robust.
  • Ensure system scalability to handle growth in web traffic and data.
  • Implement monitoring tools and set up metrics to track system health and performance.
  • Participate in on-call rotations, incident management, and postmortems.
  • Conduct performance tests to identify bottlenecks and address them.
  • Collaborate with teams to define SLOs, SLIs, and SLAs.

Skills

Go
Python
Linux
Cloud systems
Distributed systems
Monitoring
Incident management
SLIs/SLAs/SLOs

Education

Bachelor's degree in Computer Science, IT, or related field

Tools

Prometheus
Grafana
Docker
Kubernetes
Git

Job description

TikTok USDS Joint Venture invites a Site Reliability Engineer to join a team building and operating large-scale, distributed, fault-tolerant systems. You’ll design automation, monitor health, and contribute to on-call incident management across cross-functional teams.

Based in Seattle area, you’ll apply Go and Python skills, work with Linux, cloud platforms, and container tech like Docker and Kubernetes to scale services while upholding reliability and performance standards.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — Scalable Fault-Tolerant Systems
Site Reliability Engineer — Scalable Fault-Tolerant Systems

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 129,000 - 247,000
SRE: Scale & Automation for Global Infra
SRE: Scale & Automation for Global Infra

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000
Senior SRE - Compute: Scale, Automation & Reliability
Senior SRE - Compute: Scale, Automation & Reliability

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 178,000 - 342,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+1
Senior SRE: Scale, Resilience & Automation
Senior SRE: Scale, Resilience & Automation

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 187,000 - 360,000
Health insurance
401(k) with company match
Parental leave
+3
Site Reliability Engineer - Scale, Uptime & Observability
Site Reliability Engineer - Scale, Uptime & Observability

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 122,000 - 260,000
Medical, dental, and vision insurance
401(k) plan with company match
Paid parental leave
+6
Video Platform SRE: Reliability & Scale Lead
Video Platform SRE: Reliability & Scale Lead

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 177,000 - 342,000
Video Platform SRE: Scale Reliability & Automation
Video Platform SRE: Scale Reliability & Automation

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 112,725 - 177,840
SRE Tech Lead Manager — Observability, Reliability & Scale
SRE Tech Lead Manager — Observability, Reliability & Scale

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 208,000 - 438,000
Site Reliability Engineer – Scale, Automation & Uptime
Site Reliability Engineer – Scale, Automation & Uptime

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 122,000 - 260,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+3
SRE for AI Infra & Global System Reliability
SRE for AI Infra & Global System Reliability

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000