Production System Engineer Graduate (Server Management) - 2027 Start

ByteDance

San Jose (CA)

On-site

USD 76,000 - 128,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision insurance
401(k) with company match
Paid parental leave
Paid time off

Job summary

ByteDance seeks a Production Systems Engineer to join our Server Management DevOps team. You will work on Linux-based production environments, manage large-scale GPU server platforms, and build automation to improve reliability and efficiency in our data centers.

As a graduate you will collaborate with hardware, software, and infrastructure teams, learn advanced AI infrastructure practices, and contribute to end-to-end lifecycle management of servers across ByteDance's global data centers.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or related field.
  • Experience in systems engineering, infrastructure operations, DevOps, SRE, or related roles.
  • Solid understanding of Linux system administration (Debian/Ubuntu preferred) and troubleshooting.
  • Programming or scripting experience in Python, Bash, Go, or similar.
  • Understanding of OS, computer architecture, networking fundamentals, and storage systems.

Responsibilities

  • Server Infrastructure Operations: deploy, validate, monitor, maintain, and lifecycle-manage large-scale server fleets (CPU & GPU).
  • Automation Development: create scripts and tools to reduce manual work and boost efficiency.
  • Linux Systems: troubleshoot OS, hardware, storage, networking, and performance issues.
  • GPU and AI Infrastructure: work with GPU platforms and AI tooling to improve reliability.
  • Monitoring and Data Analysis: review health metrics and infrastructure data to identify trends and risks.
  • AI for Infrastructure Operations: explore applying AI/LLMs to troubleshooting and automation.

Skills

Linux admin
Python scripting
Go programming
Networking basics
Hardware knowledge

Education

Bachelor's or Master's in CS/CE/IT

Tools

Docker
Kubernetes
CUDA
REST APIs

Job description

Responsibilities

The Server Management DevOps team is responsible for the end-to-end lifecycle management of servers across our self-built data centers in the United States and Europe.

Our scope covers the complete server lifecycle, including new hardware introduction, data center delivery, production operations, hardware maintenance, configuration changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse.

The team serves as a central coordination point between multiple functions, including:

  • Hardware New Product Introduction (NPI)
  • Server and data center operations
  • Field maintenance and infrastructure management
  • Hardware vendors and service providers
  • Supply chain and asset management
  • Infrastructure platform and automation engineering teams

Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle.

We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.

Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.

Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.

We are looking for a motivated Production Systems Engineer who is passionate about Linux systems, server hardware, automation, AI infrastructure, and large-scale data center operations.

As a Production Systems Engineer, you will work alongside experienced infrastructure engineers on real production challenges involving large-scale server fleets, GPU infrastructure, automation platforms, and AI-assisted operational tools. You will contribute to building automation, improving operational efficiency, troubleshooting production issues, and supporting the lifecycle management of servers deployed across ByteDance's global data centers.

Key Responsibilities
  • Server Infrastructure Operations: Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers.
  • Automation Development: Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency.
  • Linux Systems: Work with Linux-based production environments and help troubleshoot operating system, hardware, storage, networking, and performance-related issues.
  • GPU and AI Infrastructure: Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements.
  • Monitoring and Data Analysis: Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement.
  • AI for Infrastructure Operations: Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision-making.
  • Strong analytical and troubleshooting skills with the ability to learn unfamiliar technologies quickly.
  • Good communication skills and the ability to collaborate effectively in cross-functional engineering teams.
Qualifications
Minimum Qualifications
  • Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.
  • Experience in systems engineering, infrastructure operations, DevOps, Site Reliability Engineering, or related technical roles, or equivalent hands‑on project experience.
  • Solid understanding of Linux system (Debian or Ubuntu preferred) administration and troubleshooting.
  • Programming or scripting experience in Python, Bash, Go, or another modern programming language.
  • Understanding of operating systems, computer architecture, networking fundamentals, and storage systems.
Preferred Qualifications
  • Experience developing automation tools or infrastructure software using Python, Bash, Go, or similar languages.
  • Working with server hardware, PC building, homelabs, or data center infrastructure.
  • Experience working with NVIDIA GPU platforms, AI infrastructure, CUDA, or high‑performance computing environments.
  • Applying networking fundamentals, including TCP/IP, DNS, DHCP, VLANs, and routing.
  • Hands‑on experience through internships, research, open‑source projects, homelabs, technical competitions, or personal engineering projects.
  • Experience with infrastructure monitoring, observability, logging, or telemetry platforms.
  • Experience with Git, Docker, Kubernetes, REST APIs, SQL, or infrastructure automation frameworks such as Ansible.
Job Information
【For Pay Transparency】Compensation Description (Annually)

The base salary range for this position in the selected city is $76000 - $128000 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

  • Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  • Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and
  • Exercising sound judgment.
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an “Always Day 1” mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Reasonable Accommodation

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production System Engineer Graduate (Server Management) - 2027 Start
Production System Engineer Graduate (Server Management) - 2027 Start

ByteDance • New York (NY)

On-site
USD 120,000 - 180,000
Production System Engineer Graduate (Server Management) - 2027 Start
Production System Engineer Graduate (Server Management) - 2027 Start

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Infrastructure Delivery Engineer Graduate (Server Delivery) - 2027 Start
Infrastructure Delivery Engineer Graduate (Server Delivery) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 76,000 - 128,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+1
Senior Production System Engineer - San Jose
Senior Production System Engineer - San Jose

ByteDance • San Jose (CA)

On-site
USD 115,000 - 288,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Senior Server Operations & Maintenance Engineer Technology - Infrastructure San Jose Regular
Senior Server Operations & Maintenance Engineer Technology - Infrastructure San Jose Regular

ByteDance • San Jose (CA)

On-site
USD 115,000 - 288,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Backend Development Engineer Graduate (Infrastructure Platform Delivery) - 2027 Start
Backend Development Engineer Graduate (Infrastructure Platform Delivery) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Cloud Site Reliability Engineer - DCS Cloud Seattle Regular
Cloud Site Reliability Engineer - DCS Cloud Seattle Regular

ByteDance • Seattle (WA)

On-site
USD 129,000 - 342,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+5
Site Reliability Graduate (Data Infrastructure) - 2027 Start
Site Reliability Graduate (Data Infrastructure) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 150,000 - 190,000
Software Engineer Graduate (Data-Speech-Product RD-Engineering-US) - 2027 Start
Software Engineer Graduate (Data-Speech-Product RD-Engineering-US) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Software Engineer Graduate (AI Infrastructure-Compute Efficiency & Scheduling) - 2027 Start
Software Engineer Graduate (AI Infrastructure-Compute Efficiency & Scheduling) - 2027 Start

ByteDance • Seattle (WA)

On-site
USD 122,000 - 243,000
Medical, dental, and vision insurance
401(k) with company match
Parental leave
+4