SRE Tech Lead Manager: Global Reliability & Observability

TikTok USDS Joint Venture

San Jose (CA)

On-site

USD 209,000 - 438,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

TikTok USDS Joint Venture LLC's Product Engineering SRE team builds and runs large-scale, globally distributed, fault-tolerant systems. You will own production reliability and drive observability and automation across service mesh architectures.

Responsibilities include technical leadership, architectural decisions for service mesh, incident response, and setting SLOs and error budgets while mentoring engineers and growing the team.

Qualifications

  • 5+ years designing, analyzing, and troubleshooting large-scale distributed systems.
  • Experience leading a small to mid-size team with hands-on contributions.
  • Strong Unix/Linux internals and networking fundamentals.
  • Proficiency in production-grade code in Go, Python, Java, or similar.
  • Proven track record of establishing and implementing SRE best practices.
  • Experience developing and maintaining service level objectives (SLOs) and error budgets.

Responsibilities

  • Provide technical leadership to a team of SREs focused on observable, fault-tolerant systems.
  • Drive architectural decisions for large-scale, globally distributed service mesh architectures.
  • Establish and maintain production ownership models, incident response protocols, and SLOs.
  • Develop strategic roadmaps for observability and automation initiatives that enhance system reliability.
  • Balance technical contributions with people management responsibilities, including career development, performance evaluations, and team growth.
  • Foster a culture of reliability, continuous improvement, and knowledge sharing within your team and across the organization.
  • Lead security initiatives to safeguard critical assets, partnering with security and compliance teams to implement robust protocols that ensure data protection and regulatory compliance across all services.

Skills

Distributed systems
Team leadership
Go / Python / Java
Unix/Linux internals
SRE best practices
SLOs and error budgets

Job description

TikTok USDS Joint Venture LLC's Product Engineering SRE team builds and runs large-scale, globally distributed, fault-tolerant systems. You will own production reliability and drive observability and automation across service mesh architectures.

Responsibilities include technical leadership, architectural decisions for service mesh, incident response, and setting SLOs and error budgets while mentoring engineers and growing the team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Tech Lead Manager — Observability & Reliability
SRE Tech Lead Manager — Observability & Reliability

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 209,000 - 438,000
Medical insurance
401(k) with company match
Paid parental leave
SRE Tech Lead Manager, Product Reliability
SRE Tech Lead Manager, Product Reliability

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 209,000 - 438,000
Senior SRE Lead: Observability, Scale & Reliability
Senior SRE Lead: Observability, Scale & Reliability

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 198,000 - 416,000
Senior SRE: Global Resilience & Automation
Senior SRE: Global Resilience & Automation

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 187,000 - 360,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
Video Platform SRE Lead & Architect
Video Platform SRE Lead & Architect

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 198,000 - 416,000
Medical Insurance
401(k) match
Parental leave
+2
SRE, Applied ML — Scale, Reliability & Automation
SRE, Applied ML — Scale, Reliability & Automation

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
Senior Data Infra SRE: Reliability, Scale & Automation
Senior Data Infra SRE: Reliability, Scale & Automation

TikTok • Seattle (WA)

On-site
USD 207,000 - 368,000
SRE Engineer, AI Infrastructure — Scale & Automation
SRE Engineer, AI Infrastructure — Scale & Automation

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 129,000 - 247,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Site Reliability Engineer - Scale, Uptime & Observability
Site Reliability Engineer - Scale, Uptime & Observability

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 122,574 - 259,200
Medical, dental, and vision insurance
401(k) plan with company match
Paid parental leave
+6
Site Reliability Engineer - Build resilient systems
Site Reliability Engineer - Build resilient systems

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 119,000 - 187,000
Medical/Dental/Vision insurance
401(k) with company match
Parental leave
+2