Production System Engineer Graduate (Server Management) - 2027 Start

ByteDance

New York (NY)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance in New York is seeking a Production Systems Engineer to join the Infrastructure team. You will manage end-to-end server lifecycles, from provisioning to decommissioning, across our data centers with a focus on Linux systems, automation, and AI-enabled tooling.

You will develop automation, monitor hardware and software health, and collaborate with multiple engineering groups to improve reliability and efficiency in large-scale environments.

Qualifications

  • Bachelor's or Master's degree in CS/Engineering/IT or related field.
  • Experience in systems engineering, infrastructure operations, DevOps, or SRE roles.
  • Solid Linux administration and troubleshooting (Debian/Ubuntu preferred).
  • Programming or scripting in Python, Bash, Go, or similar.

Responsibilities

  • Server Infrastructure Operations: deploy, validate, monitor, and manage large-scale server fleets.
  • Automation Development: build scripts and tools to reduce manual work and improve efficiency.
  • Linux Systems: troubleshoot OS, hardware, storage, networking, and performance issues.
  • GPU and AI Infrastructure: support GPU platforms and AI tooling for operations.
  • Monitoring and Data Analysis: analyze health metrics to find risks and improvements.
  • AI for Infrastructure Operations: explore LLM applications for automation and knowledge management.
  • Collaborate across cross-functional teams with strong communication.

Skills

Analytical skills
Troubleshooting
Communication skills
Cross-functional collaboration

Education

Bachelor's or Master's in CS/Engineering/IT

Tools

Git
Docker
Kubernetes
REST APIs
SQL
Ansible

Job description

Join us as we work together to inspire creativity and enrich life around the globe.

Location:

New York

Team:

Infrastructure

Employment Type:

Regular

Job Code:

A55465

Share this listing:

Responsibilities

The Server Management DevOps team is responsible for the end-to-end lifecycle management of servers across our self-built data centers in the United States and Europe. Our scope covers the complete server lifecycle, including new hardware introduction, data center delivery, production operations, hardware maintenance, configuration changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse.

The team serves as a central coordination point between multiple functions, including:

  • Hardware New Product Introduction (NPI)
  • Server and data center operations
  • Field maintenance and infrastructure management
  • Hardware vendors and service providers
  • Supply chain and asset management
  • Infrastructure platform and automation engineering teams

Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle.

We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.

We are looking for a motivated Production Systems Engineer who is passionate about Linux systems, server hardware, automation, AI infrastructure, and large-scale data center operations. As a Production Systems Engineer, you will work alongside experienced infrastructure engineers on real production challenges involving large-scale server fleets, GPU infrastructure, automation platforms, and AI-assisted operational tools. You will contribute to building automation, improving operational efficiency, troubleshooting production issues, and supporting the lifecycle management of servers deployed across ByteDance's global data centers.

Key Responsibilities
  • Server Infrastructure Operations: Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers.
  • Automation Development: Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency.
  • Linux Systems: Work with Linux-based production environments and help troubleshoot operating system, hardware, storage, networking, and performance-related issues.
  • GPU and AI Infrastructure: Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements.
  • Monitoring and Data Analysis: Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement.
  • AI for Infrastructure Operations: Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision-making.
  • Strong analytical and troubleshooting skills with the ability to learn unfamiliar technologies quickly.
  • Good communication skills and the ability to collaborate effectively in cross-functional engineering teams.
Minimum Qualifications
  • Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.
  • Experience in systems engineering, infrastructure operations, DevOps, Site Reliability Engineering, or related technical roles, or equivalent hands‑on project experience.
  • Solid understanding of Linux system (Debian or Ubuntu preferred) administration and troubleshooting.
  • Programming or scripting experience in Python, Bash, Go, or another modern programming language.
  • Understanding of operating systems, computer architecture, networking fundamentals, and storage systems.
Preferred Qualifications
  • Experience developing automation tools or infrastructure software using Python, Bash, Go, or similar languages.
  • Working with server hardware, PC building, homelabs, or data center infrastructure.
  • Experience working with NVIDIA GPU platforms, AI infrastructure, CUDA, or high-performance computing environments.
  • Applying networking fundamentals, including TCP/IP, DNS, DHCP, VLANs, and routing.
  • Hands‑on experience through internships, research, open‑source projects, homelabs, technical competitions, or personal engineering projects.
  • Experience with infrastructure monitoring, observability, logging, or telemetry platforms.
  • Experience with Git, Docker, Kubernetes, REST APIs, SQL, or infrastructure automation frameworks such as Ansible.
Job Information

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Reasonable Accommodation

ByteDance is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/RA-request.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production System Engineer Graduate (Server Management) - 2027 Start
Production System Engineer Graduate (Server Management) - 2027 Start

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Production System Engineer Graduate (Server Management) - 2027 Start
Production System Engineer Graduate (Server Management) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 76,000 - 128,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+1
Production System Engineer Project Intern (Server DevOps) - 2027 Start Technology - Infrastruct[...]
Production System Engineer Project Intern (Server DevOps) - 2027 Start Technology - Infrastruct[...]

Bytedance • San Jose (CA)

On-site
USD 34,000 - 62,000
Site Reliability Graduate (Data Infrastructure) - 2027 Start
Site Reliability Graduate (Data Infrastructure) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 150,000 - 190,000
Infrastructure Delivery Engineer Graduate (Server Delivery) - 2027 Start
Infrastructure Delivery Engineer Graduate (Server Delivery) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 76,000 - 128,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+1
Infrastructure Delivery Engineer Graduate (Server Delivery) - 2027 Start Technology - Infrastru[...]
Infrastructure Delivery Engineer Graduate (Server Delivery) - 2027 Start Technology - Infrastru[...]

Pangle • San Jose (CA)

On-site
USD 100,000 - 140,000
Backend and Infra Software Engineer Graduate (Dev Infra US) - 2027 Start Technology - Backend B[...]
Backend and Infra Software Engineer Graduate (Dev Infra US) - 2027 Start Technology - Backend B[...]

Bytedance • San Jose (CA)

On-site
USD 150,000 - 190,000
Senior Production System Engineer - San Jose
Senior Production System Engineer - San Jose

ByteDance • San Jose (CA)

On-site
USD 115,000 - 288,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start
Software Engineer Graduate (AML-Engine-Orchestration) - 2027 Start

ByteDance • Seattle (WA)

On-site
USD 110,000 - 150,000
Backend Development Engineer Graduate (Infrastructure Platform Delivery) - 2027 Start
Backend Development Engineer Graduate (Infrastructure Platform Delivery) - 2027 Start

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical Insurance
Dental Insurance
Vision Insurance
+2