Senior Site Reliability Engineer - Data Infrastructure (San Jose)

ByteDance

San Jose (CA)

On-site

USD 212,800 - 387,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a Senior Site Reliability Engineer for Data Infrastructure in San Jose to keep our large-scale data systems reliable and efficient. You will work on hands-on operations, respond to alerts, manage production changes, and automate routine tasks alongside senior engineers.

The role emphasizes reliability, scale, and cost optimization, with rotations on on-call, collaboration across time zones, and involvement in data center and AI infrastructure efforts.

Qualifications

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • 5+ years of experience in a Site Reliability Engineering, Production Engineering, or similar role.
  • Strong proficiency in a programming or scripting language (e.g., Go, Python, Bash) for automation and tool development.
  • Deep understanding of Linux/Unix operating systems, networking fundamentals (TCP/IP, DNS), and distributed systems.

Responsibilities

  • Incident response and postmortems for critical production issues with blameless reviews.
  • Define and maintain SLOs for critical data services. Manage error budgets to balance reliability work with features.
  • Lead capacity planning, performance tuning, and resource management to stay within budget.
  • Design automation and AI orchestration to reduce toil and improve safety and efficiency.
  • Uphold production operations standards including runbooks, monitoring, and change management for deployments.
  • Lead data center and AI infrastructure initiatives to ensure high availability and peak performance.
  • Mentor junior SREs and collaborate with other teams across time zones.

Skills

Go
Python
Bash
Linux/Unix
Networking basics
Distributed systems

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
MySQL
Redis
Kafka
Flink

Job description

Senior Site Reliability Engineer - Data Infrastructure (San Jose)

Location: San Jose

Team: Infrastructure

Employment Type: Regular

Responsibilities

The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power our products. We manage a massive, distributed environment built on technologies like Kubernetes, Redis, MySQL, and Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform that all product teams depend on. We are the guardians of production, ensuring our data systems run smoothly, nonstop. This role includes participation in a rotational on‑call schedule to ensure nonstop coverage for our critical data infrastructure. You will be expected to respond to, troubleshoot, and resolve production incidents. Our team collaborates across multiple time zones, and you will engage in rigorous change management and post‑incident review processes to maintain system stability.

Role Summary: As a Site Reliability Engineer, you will be on the front lines of keeping our large‑scale data systems running reliably and efficiently. You will focus on hands‑on operational work, from responding to alerts and managing production changes to automating routine tasks. This role is an excellent opportunity to develop deep expertise in modern infrastructure technologies and SRE practices while working alongside senior engineers to solve challenging problems.

  • Incident response and postmortems: Act as an incident commander for critical production issues, guiding the team through triage and resolution. Drive deep, blameless post‑incident reviews and ensure that follow‑up actions are implemented to prevent recurrence.
  • SLO/SLA and error budgets: Define, negotiate, and maintain Service Level Objectives (SLOs) for critical data services. Champion the use of error budgets to balance reliability work with feature development.
  • Capacity and cost optimization: Lead initiatives in capacity planning, performance tuning, and resource management. Develop strategies and automation to ensure our infrastructure scales efficiently and stays within budget.
  • Pragmatic automation and AI orchestration: Design and build automation and leverage AI Agents to eliminate operational toil, improve deployment safety, and enhance overall operational efficiency. Focus on creating maintainable, robust tools and intelligent workflows that make the entire team more effective.
  • Operational excellence and change management: Uphold and improve our standards for production operations, including runbooks, monitoring, and alerting. Vet complex changes and deployments to ensure they meet our bar for production readiness.
  • Data Center and AI Infrastructure: Lead the construction, maintenance, and optimization of data centers and specialized AI infrastructure, ensuring high availability and peak performance for complex AI‑driven workloads.
  • Cross‑team influence and mentorship: Act as a subject matter expert on reliability, consulting with application development and other infrastructure teams. Mentor junior SREs, helping them develop their technical and operational skills.
Qualifications

Minimum Qualifications:

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • 5+ years of experience in a Site Reliability Engineering, Production Engineering, or similar role.
  • Strong proficiency in a programming or scripting language (e.g., Go, Python, Bash) for automation and tool development.
  • Deep understanding of Linux/Unix operating systems, networking fundamentals (TCP/IP, DNS), and distributed systems.

Preferred Qualifications:

  • Extensive hands‑on experience managing large‑scale data infrastructure (e.g., MySQL, Redis, Kafka, Flink).
  • Proven experience with container orchestration technologies, particularly Kubernetes, in a production environment.
  • Expertise in designing, analyzing, and troubleshooting large‑scale distributed systems.
  • A systematic problem‑solving approach, coupled with strong communication skills and a sense of ownership.
  • Experience leading incident response for complex, high‑impact events.
  • Experience in the operation and construction of Data Centers is a big plus.
Job Information

The base salary range for this position in the selected city is $212,800 - $387,600 annually.

Benefits

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short‑term and long‑term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

Reasonable Accommodation

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead - Data Infrastructure Site Reliability
Tech Lead - Data Infrastructure Site Reliability

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500
Medical insurance
Dental and vision insurance
401(k) with company match
+7
Senior SRE - Data Infrastructure & Reliability
Senior SRE - Data Infrastructure & Reliability

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
Tech Lead Cloud Site Reliability Engineer - DCS Cloud

Socket.dev • Seattle (WA)

On-site
USD 232,560 - 427,500
Site Reliability Engineer - Data San Jose Regular
Site Reliability Engineer - Data San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 136,000 - 360,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+3
Senior Software Engineer, Cloud Infrastructure
Senior Software Engineer, Cloud Infrastructure

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical, dental and vision insurance
401(k) with company match
Parental leave
+6
Site Reliability Engineer - Data (Seattle) Seattle Regular
Site Reliability Engineer - Data (Seattle) Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 177,000 - 342,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Tech Lead - Data Infrastructure Site Reliability
Tech Lead - Data Infrastructure Site Reliability

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Health insurance
401(k) with company match
Paid parental leave
+3
Senior Software Engineer - Compute Infrastructure (Orchestration & Scheduling)
Senior Software Engineer - Compute Infrastructure (Orchestration & Scheduling)

ByteDance • San Jose (CA)

On-site
USD 156,000 - 387,600
Medical insurance
401(k) plan with company match
Parental leave
+6
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
Tech Lead Cloud Site Reliability Engineer - DCS Cloud

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Site Reliability Engineer Intern (Data Infra) - 2027 Fall
Site Reliability Engineer Intern (Data Infra) - 2027 Fall

ByteDance • Seattle (WA)

On-site
USD 36,000 - 59,000
Health insurance
Life insurance
Wellbeing benefits
+2