Senior Site Reliability Engineer (Multiple Positions)

ByteDance

Seattle (WA)

On-site

USD 212,000 - 368,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance in Bellevue, WA, is seeking a Site Reliability Engineer to ensure the highest availability of large-scale, fault-tolerant systems. You will design and deploy automation to scale reliability across backend and cloud-native services.

You will measure latency, health, and incident response, performing root cause analysis to guide future product design. Mentoring junior engineers and interns will be part of the role.

Qualifications

  • Master's degree or foreign equivalent in Computer Science, Engineering, Information Systems, Data Science, Mathematics, or a related field; 2 years of related work experience OR Bachelor's degree with 5 years of post-bachelor’s related work experience.
  • 2 years of required experience in: reliability support for critical components; SDLC across backend/cloud native projects; runbooks for alerts/ troubleshooting; data services operations and SLA management; Linux admin including virtualization and containers.

Responsibilities

  • Provide site reliability engineering support to ensure highest level of availability of large-scale, fault tolerant systems.
  • Deliver tools and software to improve reliability, scalability and operability, including designing, developing and deploying automation to scale with quality.
  • Measure and monitor availability, latency and overall service health.
  • Practice sustainable incident response and postmortems, performing root cause analysis of incidents to influence future product design and response activities.
  • Establish solid design and best practices for engineers as well as non-technical team members.
  • Mentor junior engineers and interns.

Skills

SRE expertise
Linux administration
SDLC
Incident response
Runbooks
SLA management
Monitoring & observability

Education

Master's degree in CS/Engineering/IS/Data Science/Math
Bachelor's degree in CS/Engineering/IS/Data Science/Math

Job description

Join us as we work together to inspire creativity and enrich life around the globe.

Location:

Seattle

Team:

Technology

Employment Type:

Regular

Job Code:

A72769A

Share this listing:
Responsibilities
  • Provide site reliability engineering support to ensure highest level of availability of large-scale, fault tolerant systems.
  • Deliver tools and software to improve the reliability, scalability and operability of services, including designing, developing and deploying automation to sustainably scale with quality.
  • Measure and monitor availability, latency and overall service health.
  • Practice sustainable incident response and postmortems, performing root cause analysis of incidents to influence future product design and response activities.
  • Establish solid design and best practices for engineers as well as non-technical team members.
  • Mentor junior engineers and intern.
About ByteDance

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join Us

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless.Join us.

About the Team

Our team plays a crucial role in ensuring the company’s success. We seek people who are willing to learn and put in the effort to solve problems. Our challenges are not your regular day-to-day problems - you’ll be part of a team that’s developing new solutions to new challenges. It’s working fast, at scale, and we’re making a difference. We are looking for talents to join us on this exciting journey!

Qualifications
  • Must have a Master's degree or foreign equivalent degree in Computer Science, Engineering (any), Information Systems, Data Science, Mathematics, or a related field, and 2 years of related work experience; OR a Bachelor's degree or foreign equivalent degree in Computer Science, Engineering (any), Information Systems, Data Science, Mathematics, or a related field, and 5 years of post-bachelor’s, progressive related work experience.
  • Of the required experience, must have 2 years of experience in each of the following:Providing functionality and reliability support for critical site components by measuring and monitoring availability, latency, and overall system health; Working across all phases of the SDLC, including requirements gathering and analysis, design, development, implementation, testing, deployment, and maintenance of back-end and cloud native projects; Creating and maintaining clear runbook instructions for services to use for alerts, troubleshooting and resolution; Coordinating and monitoring data services operations, including SLA management and system deployment; and Performing Linux administration, including monitoring performance, debugging issues, monitoring network behavior, and troubleshooting networked applications using: OS networking protocol stack and the following OS concepts: virtualization and containerization.
  • Travel Requirement: International and domestic travel required up to 10%.
Job Information

Type: Full time, 40 hours/weekLocation: Bellevue, WASalary Range: $212202 - $368220 per year

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace.

At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request#IND-DNI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Data (Seattle) Seattle Regular
Site Reliability Engineer - Data (Seattle) Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 177,000 - 342,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Cloud Site Reliability Engineer - DCS Cloud Seattle Regular
Cloud Site Reliability Engineer - DCS Cloud Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 129,000 - 342,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+5
Site Reliability Engineer - Data San Jose Regular
Site Reliability Engineer - Data San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 136,000 - 360,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+3
Solutions Engineer (Multiple Positions)
Solutions Engineer (Multiple Positions)

ByteDance • New York (NY)

On-site
USD 156,000 - 317,000
Site Reliability Graduate (Data Infrastructure) - 2027 Start
Site Reliability Graduate (Data Infrastructure) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 150,000 - 190,000
Cloud Site Reliability Engineer - DCS Cloud San Jose Regular
Cloud Site Reliability Engineer - DCS Cloud San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 136,000 - 360,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid personal time off
+1
Backend and Infra Software Engineer Graduate (Dev Infra US) - 2027 Start Technology - Backend B[...]
Backend and Infra Software Engineer Graduate (Dev Infra US) - 2027 Start Technology - Backend B[...]

Bytedance • San Jose (CA)

On-site
USD 150,000 - 190,000
Production System Engineer Graduate (Server Management) - 2027 Start
Production System Engineer Graduate (Server Management) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 76,000 - 128,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+1
Software Engineer Graduate (Data-Speech-Product RD-Engineering-US) - 2027 Start
Software Engineer Graduate (Data-Speech-Product RD-Engineering-US) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Software Engineer Intern (Distributed NoSQL Database Systems) - 2027 Summer
Software Engineer Intern (Distributed NoSQL Database Systems) - 2027 Summer

ByteDance • Seattle (WA)

On-site
USD 30,000 - 40,000