Tech Lead - Data Infrastructure Site Reliability

ByteDance

San Jose (CA)

On-site

USD 244,800 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) with company match
Paid parental leave
Disability coverage
Wellbeing benefits
Paid personal time

Job summary

ByteDance in San Jose seeks a Tech Lead for Data Infrastructure Site Reliability to steer the design, development, and operation of large-scale cloud infrastructure. You will collaborate with Advertising, ML, E-commerce, and Core Infra teams to ensure reliability and performance across platforms.

The role emphasizes hands-on leadership, deep technical expertise, and the ability to influence across organizations while driving automation and cost-efficient solutions.

Qualifications

  • 5+ years of experience in Site Reliability Engineering, software development, or related fields.
  • Hands-on expertise in databases, Kubernetes, or big data systems.
  • Strong knowledge of system architecture, distributed systems, and collaboration.

Responsibilities

  • Hands-on design, development, and operation of large-scale cloud infrastructure.
  • Collaborate with cross-functional teams to improve reliability, performance, and scalability.
  • Lead automation initiatives to reduce toil and improve efficiency.
  • Troubleshoot complex production issues and drive long-term reliability.
  • Promote best practices in observability, performance, and cost efficiency.
  • Communicate complex technical concepts to technical and non-technical stakeholders.

Skills

SRE experience
Cloud-based systems
Databases (SQL/NoSQL)
Kubernetes
Big Data

Job description

Tech Lead - Data Infrastructure Site Reliability

Location: San Jose

Team: Technology

Employment Type: Regular

Job Code: A246362

Team Introduction: Our Site Reliability Engineering (SRE) team blends software and systems engineering to build and operate large-scale data infrastructure with high reliability and efficiency. We provide a dependable cloud environment that powers our global business. In this role, you will leverage your expertise in data center architecture, data infrastructure services, and systems and tools development to solve complex scaling and reliability challenges. We’re looking for a Technical Lead (SRE) who can provide deep technical leadership, drive architectural improvements, and collaborate effectively across multiple organizations. You’ll partner with engineering, product, data, and infrastructure teams to deliver resilient, scalable platforms. This is a highly technical, hands‑on role that requires strong problem‑solving ability, clear communication, and the ability to influence without formal authority.

Responsibilities
  • Strong hands‑on skills in the design, development, and operation of large‑scale cloud infrastructure and distributed systems.
  • Collaborate with cross‑functional teams (Advertising, Machine Learning, E‑commerce, and Core Infra) to drive system reliability, performance, and scalability.
  • Lead initiatives to automate operations, eliminate toil, and improve overall system efficiency.
  • Troubleshoot complex production issues, perform root‑cause analysis, and drive long‑term reliability improvements.
  • Promote best practices in system design, observability, performance optimization, and cost efficiency.
  • Communicate complex technical concepts effectively to both technical and non‑technical stakeholders.
Qualifications
Minimum Qualifications
  • 5+ years of experience in Site Reliability Engineering, Software Development, or related fields, with a strong focus on designing, building, scaling, and operating cloud‑based systems.
  • Deep hands‑on expertise in at least one of the following areas:
    • Databases (SQL/NoSQL)
    • Kubernetes or container orchestration
    • Big Data processing and storage systems (streaming and batch)
  • Strong knowledge of system architecture, distributed systems, and performance bottlenecks.
  • Excellent communication and collaboration skills, with experience working across engineering, product, and data science teams.
Preferred Qualifications
  • Proven track record of driving automation, tooling, and process improvements that enhance reliability and efficiency.
  • Experience in cost optimization and performance tuning at scale, backed by data‑driven decision making.
  • Thought leadership in adopting new technologies, improving operational practices, and influencing system design.
Job Information

The base salary range for this position in the selected city is $244,800 – $450,000 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day‑one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short‑term and long‑term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

Equal Employment Opportunity Statement

For Los Angeles County (unincorporated) Candidates:

  • Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:
  • Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  • Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems;
  • Exercising sound judgment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead - Data Infrastructure Site Reliability
Tech Lead - Data Infrastructure Site Reliability

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500
Medical insurance
Dental and vision insurance
401(k) with company match
+7
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
Senior Site Reliability Engineer - Data Infrastructure (San Jose)

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Tech Lead - Machine Learning Platform Engineer
Tech Lead - Machine Learning Platform Engineer

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+2
Software Engineering Manager II, Site Reliability Engineering
Software Engineering Manager II, Site Reliability Engineering

Google • Los Angeles (CA)

On-site
USD 197,000 - 291,000
Competitive salary
Bonus
Equity
+1
Senior Software Engineer, Cloud Infrastructure
Senior Software Engineer, Cloud Infrastructure

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical, dental and vision insurance
401(k) with company match
Parental leave
+6
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NAB Leadership Foundation • San Antonio (TX)

On-site
USD 90,000 - 120,000
Employer sponsored medical, dental, and vision coverage
Paid vacation and sick time
401K plan
+1
Tech Lead Software Engineer, Programming Language
Tech Lead Software Engineer, Programming Language

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Medical Insurance
401k Matching
Parental Leave
+6
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Illumio • San Jose (CA)

On-site
USD 120,000 - 150,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Green Dot Corporation • Town of Florida (NY), Los Angeles (CA)

Hybrid
USD 139,000 - 200,000