Senior SRE - Data Infrastructure & Reliability

ByteDance

San Jose (CA)

On-site

USD 212,800 - 387,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a Senior Site Reliability Engineer for Data Infrastructure in San Jose to keep our large-scale data systems reliable and efficient. You will work on hands-on operations, respond to alerts, manage production changes, and automate routine tasks alongside senior engineers.

The role emphasizes reliability, scale, and cost optimization, with rotations on on-call, collaboration across time zones, and involvement in data center and AI infrastructure efforts.

Qualifications

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • 5+ years of experience in a Site Reliability Engineering, Production Engineering, or similar role.
  • Strong proficiency in a programming or scripting language (e.g., Go, Python, Bash) for automation and tool development.
  • Deep understanding of Linux/Unix operating systems, networking fundamentals (TCP/IP, DNS), and distributed systems.

Responsibilities

  • Incident response and postmortems for critical production issues with blameless reviews.
  • Define and maintain SLOs for critical data services. Manage error budgets to balance reliability work with features.
  • Lead capacity planning, performance tuning, and resource management to stay within budget.
  • Design automation and AI orchestration to reduce toil and improve safety and efficiency.
  • Uphold production operations standards including runbooks, monitoring, and change management for deployments.
  • Lead data center and AI infrastructure initiatives to ensure high availability and peak performance.
  • Mentor junior SREs and collaborate with other teams across time zones.

Skills

Go
Python
Bash
Linux/Unix
Networking basics
Distributed systems

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
MySQL
Redis
Kafka
Flink

Job description

ByteDance is seeking a Senior Site Reliability Engineer for Data Infrastructure in San Jose to keep our large-scale data systems reliable and efficient. You will work on hands-on operations, respond to alerts, manage production changes, and automate routine tasks alongside senior engineers.

The role emphasizes reliability, scale, and cost optimization, with rotations on on-call, collaboration across time zones, and involvement in data center and AI infrastructure efforts.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Graduate: Data Platform Reliability & Automation
SRE Graduate: Data Platform Reliability & Automation

ByteDance • San Jose (CA)

On-site
USD 150,000 - 190,000
Tech Lead, Data Infrastructure & SRE (Cloud-Scale)
Tech Lead, Data Infrastructure & SRE (Cloud-Scale)

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Health insurance
401(k) with company match
Paid parental leave
+3
Senior Site Reliability Engineer - Data Infrastructure (San Jose)
Senior Site Reliability Engineer - Data Infrastructure (San Jose)

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Global Data Center SRE | Reliable Infrastructure
Global Data Center SRE | Reliable Infrastructure

ByteDance • San Jose (CA)

On-site
USD 210,000 - 330,000
Tech Lead: Data Infrastructure & Site Reliability
Tech Lead: Data Infrastructure & Site Reliability

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500
Medical insurance
Dental and vision insurance
401(k) with company match
+7
Senior Backend & Infra Engineer, Data Platforms
Senior Backend & Infra Engineer, Data Platforms

ByteDance • San Jose (CA)

On-site
USD 156,000 - 387,600
Cloud SRE Tech Lead - Global Infra & Automation
Cloud SRE Tech Lead - Global Infra & Automation

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Senior SRE: Global Traffic & Edge Infrastructure
Senior SRE: Global Traffic & Edge Infrastructure

Bytedance • San Jose (CA)

On-site
USD 140,000 - 190,000
Senior Data Infrastructure SRE: Reliability & Scale
Senior Data Infrastructure SRE: Reliability & Scale

TikTok • Seattle (WA)

On-site
USD 202,000 - 368,000
Data Platform SRE: Scale Reliable Big Data Systems
Data Platform SRE: Scale Reliable Big Data Systems

TikTok • San Jose (CA)

Hybrid
USD 118,000 - 260,000