Production System Engineer

ByteDance

Greater London

On-site

GBP 70,000 - 120,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a hands-on Production Systems Engineer to manage Linux-based server infrastructure, hardware lifecycle, and automation in global data centers. You will work on deployment, monitoring, and maintenance of large-scale fleets, including GPUs, while collaborating with multiple engineering and operations teams to enhance reliability and efficiency.

The role welcomes engineers across levels, with ownership expanding from hands-on engineering to leading complex global initiatives.

Qualifications

  • Bachelor's degree or above in a technical field (CS/CE/EE/IT).
  • 2 years of experience in systems engineering or DevOps/SRE roles or equivalent hands-on work.
  • Strong Linux administration and troubleshooting foundation.
  • Programming or scripting in Python, Bash, Go, or similar for tooling/automation.
  • Hands-on experience troubleshooting Linux-based systems, hardware, storage, networking, and performance.
  • Sharp analytical and problem-solving skills; quick to learn new tech.
  • Good communication and collaboration across teams and regions.

Responsibilities

  • Manage deployment, validation, monitoring, maintenance, and lifecycle of large server fleets (CPU/GPU).
  • Develop automation scripts and tools to reduce manual workload.
  • Troubleshoot Linux systems, hardware, storage, networking, and performance issues.
  • Gain experience with AI infrastructure and GPU server platforms; improve tooling and reliability.
  • Analyze metrics and hardware data to identify trends and risks; document processes.

Skills

Linux administration
Python
Go
Bash
SRE/DevOps
Hardware troubleshooting
Communication

Education

Bachelor's degree in CS/CE/EE/IT

Tools

Docker
Kubernetes
Redfish
Firmware
BMC
PCIe/NVMe

Job description

Responsibilities

About the Team

The Server Management DevOps team is responsible for the end-to-end lifecycle management of servers across ByteDance’s self-built data centers in the United States and Europe.

Our scope covers new hardware introduction, data center delivery, production operations, hardware maintenance, configuration and firmware changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse.

The team serves as a central coordination point between multiple functions, including:

  • Hardware New Product Introduction (NPI)
  • Server and data center operations
  • Field maintenance and infrastructure management
  • Hardware vendors and service providers
  • Supply chain and asset management
  • Infrastructure platform and automation engineering teams

Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle.

Role Overview

We are looking for a hands‑on Production Systems Engineer with a strong foundation in Linux systems, server infrastructure, automation, and production operations. This role is open to engineers across a range of experience levels, from early-career engineers with strong technical fundamentals to experienced infrastructure engineers who can take ownership of complex systems and large-scale initiatives.

The scope and level of ownership will grow with experience, ranging from hands‑on infrastructure engineering and automation development to leading complex global infrastructure initiatives across organizational boundaries.

Responsibilities
  • Server Infrastructure Operations: Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers.
  • Automation Development: Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency.
  • Linux Systems: Work with Linux‑based production environments and help troubleshoot operating system, hardware, storage, networking, and performance‑related issues.
  • GPU and AI Infrastructure: Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements. Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision‑making.
  • Monitoring and Data Analysis: Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement.
  • Technical Documentation: Create and improve technical documentation, standard operating procedures, troubleshooting guides, and internal knowledge bases.
  • Cross‑functional Collaboration: Work with infrastructure engineers, hardware teams, data center operations, platform developers, supply chain teams, and other stakeholders on global infrastructure projects.
Qualifications
Minimum Qualification(s)
  • Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.
  • 2 years of experience in systems engineering, infrastructure operations, DevOps, Site Reliability Engineering, or related technical roles, or equivalent hands‑on project experience.
  • Strong foundation in Linux system administration and troubleshooting, with an understanding of basic server architecture, operating systems, storage, networking, and hardware management concepts.
  • Programming or scripting experience in Python, Bash, Go, or another modern programming language, with the ability to develop tools or automation for infrastructure or operational tasks.
  • Hands‑on experience troubleshooting system, hardware, storage, networking, or performance‑related issues in Linux‑based environments.
  • Strong analytical and problem‑solving skills, with the ability to learn unfamiliar technologies quickly and investigate complex technical issues in a structured manner.
  • Good communication and collaboration skills, with the ability to work effectively with engineers and cross‑functional stakeholders across different technical domains and regions.
Preferred Qualification(s)
  • Familiarity with technologies such as BIOS/UEFI, BMC, firmware, PCIe, NVMe, NICs, or hardware telemetry.
  • Proficiency in Python, Go, Bash, or another programming language for production‑grade infrastructure automation, including experience designing, building, or maintaining tools and platforms used in large‑scale production environments.
  • Deep knowledge of Linux administration and troubleshooting, preferably Debian or Ubuntu, combined with strong understanding of server architecture and management technologies such as kernels, drivers, BIOS/UEFI, BMC, Redfish, firmware, PCIe, NVMe, NICs, DPUs, hardware telemetry, and failure diagnostics.
  • Proven hands‑on experience introducing and productionizing large‑scale GPU infrastructure, including ownership of hardware NPI or fleet onboarding across qualification, system integration, deployment, production validation, operational handoff, and post‑launch reliability.
  • Strong understanding of distributed AI workload behavior and performance analysis, including collective communication, multi‑node training, inference serving, GPU scheduling, checkpointing, workload‑related bottlenecks, DCGM, NCCL testing, CUDA profiling, and network‑fabric telemetry.
  • Experience building and operating monitoring, telemetry, hardware management, or automated remediation platforms at substantial scale, with measurable improvements in fleet availability, deployment efficiency, incident reduction, operational efficiency, or reliability.
  • Experience working directly with OEMs, ODMs, component suppliers, or GPU platform vendors throughout qualification, technical escalation, root‑cause analysis, and corrective‑action processes.
  • Experience with one or more advanced infrastructure technologies or engineering areas, such as containerisation and orchestration (e.g., Docker, Kubernetes), infrastructure automation frameworks (e.g., Ansible), AI‑powered automation, AI agents, Large Language Models, Retrieval‑Augmented Generation (RAG), open‑source infrastructure projects, technical publications, patents, or relevant industry standards.
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production System Engineer London Regular
Production System Engineer London Regular

ByteDance • Greater London

On-site
GBP 70,000 - 90,000
Diverse and inclusive workplace
Innovative projects
Opportunity for growth
System Software Architect - OS and Kernel Direction
System Software Architect - OS and Kernel Direction

ByteDance • Greater London

On-site
GBP 110,000 - 150,000
Edge Network Engineer - EDGE Networking
Edge Network Engineer - EDGE Networking

ByteDance • Greater London

On-site
GBP 90,000 - 140,000
Solutions Architect, Byteplus
Solutions Architect, Byteplus

ByteDance • Greater London

On-site
GBP 70,000 - 90,000
AI solution Architect, BytePlus
AI solution Architect, BytePlus

ByteDance • Greater London

On-site
GBP 90,000 - 140,000
Solutions Architect, Byteplus
Solutions Architect, Byteplus

ByteDance • Greater London

On-site
GBP 70,000 - 100,000
Forward Deployed Engineer - BytePlus
Forward Deployed Engineer - BytePlus

ByteDance • Greater London

On-site
GBP 90,000 - 140,000
Solutions Architect, Global Partnerships - Byteplus
Solutions Architect, Global Partnerships - Byteplus

ByteDance • Greater London

On-site
GBP 70,000 - 90,000
Country Lead, BytePlus
Country Lead, BytePlus

ByteDance • Greater London

On-site
GBP 120,000 - 200,000
Data Center Technical Project Manager, Data Center Development
Data Center Technical Project Manager, Data Center Development

ByteDance • Greater London

On-site
GBP 70,000 - 110,000