Site Reliability Graduate (Data Infrastructure) - 2027 Start

ByteDance

San Jose (CA)

On-site

USD 150,000 - 190,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance, San Jose, seeks a Data Infrastructure SRE to ensure reliability and efficiency of core data services powering its products. You’ll work with Kubernetes, Redis, MySQL, and Message Queue, focusing on resilience over feature development.

Expect a rotational on-call schedule to keep production healthy and incident-ready across time zones. Responsibilities include incident response, operational tasks with change control, and driving automation through scripting and AI augmentation to

Qualifications

  • Bachelor’s degree in Computer Science, Computing Engineering or related field.
  • Experience with at least one scripting language (Python, Bash, Go).
  • Solid understanding of Linux OS and networking concepts.

Responsibilities

  • Incident response and triage for production alerts and incidents.
  • Deployments, configuration changes, and system maintenance following change control processes.
  • Identify and automate repetitive tasks using scripting and AI Agents to reduce toil.

Skills

Scripting languages (Python, Bash, Go)
Linux fundamentals
Networking basics

Education

Bachelor’s degree in CS/Computing Engineering
Master’s degree (preferred)

Tools

Docker
Kubernetes
Redis
MySQL
Prometheus/Grafana/ELK

Job description

Join us as we work together to inspire creativity and enrich life around the globe.

Location:

San Jose

Team:

Technology

Employment Type:

Regular

Job Code:

A160183

Share this listing:

The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power our products. We manage a massive, distributed environment built on technologies like Kubernetes, Redis, MySQL, and Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform that all product teams depend on. We are the guardians of production, ensuring our data systems run smoothly, constantly. This role includes participation in a rotational on-call schedule to ensure constant coverage for our critical data infrastructure. You will be expected to respond to, troubleshoot, and resolve production incidents. Our team collaborates across multiple time zones, and you will engage in rigorous change management and post-incident review processes to maintain system stability.

We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume. Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.

Responsibilities
  • Incident response and triage: Serve as a first responder for production alerts and incidents. Execute established runbooks to mitigate issues and resolve production incidents. Our team collaborates across multiple time zones, and you will engage in rigorous change management and post-incident review processes to maintain system stability.
  • Operational excellence and change management: Perform routine operational tasks, such as deployments, configuration changes, and system maintenance, following established change control processes to minimize production risk.
  • Iterative automation and AI augmentation: Identify and automate repetitive manual tasks using scripting (e.g., Python, Go, Bash) and AI Agents to reduce toil, improve operational consistency, and boost overall productivity.
  • Observability and monitoring: Improve our observability posture by refining monitoring dashboards, tuning alert thresholds, and ensuring that our systems are sufficiently instrumented to detect and diagnose problems.
  • Data Center and AI Infrastructure: Support the daily operations, construction, and maintenance of data center environments and AI infrastructure to meet the demands of large-scale data processing.
Qualifications
Minimum Qualifications
  • Individuals who are completing or have recently completed a Bachelor's degree in Computer Science, Computing Engineering or a related discipline.
  • Experience with at least one scripting language (e.g., Python, Bash, Go).
  • Solid understanding of Linux operating systems and networking concepts.
Preferred Qualifications
  • Individuals who are completing or have recently completed a Master's degree in Computer Science, Computing Engineering or a related discipline.
  • Familiarity with container technologies like Docker and Kubernetes.
  • Hands-on experience with at least one common data store (e.g., MySQL, Redis, PostgreSQL).
  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK Stack).
  • A strong desire to learn, a proactive attitude toward problem-solving, and excellent communication skills.
  • Experience in the operation and construction of Data Centers is a big plus.
Job Information
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Reasonable Accommodation

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend and Infra Software Engineer Graduate (Dev Infra US) - 2027 Start Technology - Backend B[...]
Backend and Infra Software Engineer Graduate (Dev Infra US) - 2027 Start Technology - Backend B[...]

Bytedance • San Jose (CA)

On-site
USD 150,000 - 190,000
Site Reliability Engineer - Data San Jose Regular
Site Reliability Engineer - Data San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 136,000 - 360,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+3
Site Reliability Engineer - Data (Seattle) Seattle Regular
Site Reliability Engineer - Data (Seattle) Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 177,000 - 342,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Production System Engineer Graduate (Server Management) - 2027 Start
Production System Engineer Graduate (Server Management) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 76,000 - 128,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+1
Software Engineer Graduate (Cloud Native Infrastructure)- 2026 Start (PHD)
Software Engineer Graduate (Cloud Native Infrastructure)- 2026 Start (PHD)

ByteDance • San Jose (CA)

On-site
USD 122,574 - 316,800
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start (PhD)

ByteDance • Seattle (WA)

On-site
USD 140,000 - 210,000
Production System Engineer Graduate (Server Management) - 2027 Start
Production System Engineer Graduate (Server Management) - 2027 Start

ByteDance • New York (NY)

On-site
USD 120,000 - 180,000
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start

ByteDance • Seattle (WA)

On-site
USD 110,000 - 150,000
Backend Development Engineer Graduate (Infrastructure Platform Delivery) - 2027 Start
Backend Development Engineer Graduate (Infrastructure Platform Delivery) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Backend Software Engineer Graduate (Platform) - 2027 Start
Backend Software Engineer Graduate (Platform) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+6