Production System Engineer

ByteDance

Greater London

On-site

GBP 70,000 - 110,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

ByteDance is seeking a hands-on Production Systems Engineer to own Linux-based server infrastructure, automation, and production operations. You will contribute across large-scale deployments, performance tuning, and cross-functional projects in a global data-center environment.

The role suits engineers from early-career to senior levels and offers growth in complex, cross-region infrastructure initiatives. Strong scripting and problem-solving are essential.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, IT, or a related field.
  • 2 years in systems engineering, infra operations, DevOps, SRE, or related roles.
  • Strong Linux administration and troubleshooting fundamentals.
  • Programming or scripting in Python, Bash, Go, or similar for automation.
  • Hands-on troubleshooting in Linux-based environments (system/hardware).
  • Strong analytical and problem-solving skills with quick learnability.

Responsibilities

  • Deploy, validate, monitor, and maintain large-scale server fleets (CPU/GPU).
  • Develop scripts and automation to reduce manual work.
  • Troubleshoot Linux-based production environments (OS, hardware, storage, networking).
  • Gain exposure to AI infrastructure and GPU server platforms; contribute tooling and monitoring.
  • Analyze server health and operational metrics to identify risks and improvements.
  • Create and improve technical documentation and SOPs.
  • Collaborate with cross-functional teams on global infrastructure projects.

Skills

Linux
Scripting
Python
Go
Troubleshooting
Communication

Education

Bachelor's degree in CS/Engineering/IT

Tools

Docker
Kubernetes
Redfish
BIOS/UEFI
Firmware

Job description

Responsibilities

About the Team

The Server Management DevOps team is responsible for the end-to-end lifecycle management of servers across ByteDance’s self-built data centers in the United States and Europe. Our scope covers new hardware introduction, data center delivery, production operations, hardware maintenance, configuration and firmware changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse. The team serves as a central coordination point between multiple functions, including:

  • Hardware New Product Introduction (NPI)
  • Server and data center operations
  • Field maintenance and infrastructure management
  • Hardware vendors and service providers
  • Supply chain and asset management
  • Infrastructure platform and automation engineering teams

Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle.

Role Overview

We are looking for a hands‑on Production Systems Engineer with a strong foundation in Linux systems, server infrastructure, automation, and production operations. This role is open to engineers across a range of experience levels, from early-career engineers with strong technical fundamentals to experienced infrastructure engineers who can take ownership of complex systems and large-scale initiatives.

The scope and level of ownership will grow with experience, ranging from hands‑on infrastructure engineering and automation development to leading complex global infrastructure initiatives across organizational boundaries.

Responsibilities
  • Server Infrastructure Operations: Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers.
  • Automation Development: Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency.
  • Linux Systems: Work with Linux-based production environments and help troubleshoot operating system, hardware, storage, networking, and performance-related issues.
  • GPU and AI Infrastructure: Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements. Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision-making.
  • Monitoring and Data Analysis: Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement.
  • Technical Documentation: Create and improve technical documentation, standard operating procedures, troubleshooting guides, and internal knowledge bases.
  • Cross-functional Collaboration: Work with infrastructure engineers, hardware teams, data center operations, platform developers, supply chain teams, and other stakeholders on global infrastructure projects.
Qualifications
Minimum Qualification(s)
  • Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.
  • 2 years of experience in systems engineering, infrastructure operations, DevOps, Site Reliability Engineering, or related technical roles, or equivalent hands‑on project experience.
  • Strong foundation in Linux system administration and troubleshooting, with an understanding of basic server architecture, operating systems, storage, networking, and hardware management concepts.
  • Programming or scripting experience in Python, Bash, Go, or another modern programming language, with the ability to develop tools or automation for infrastructure or operational tasks.
  • Hands‑on experience troubleshooting system, hardware, storage, networking, or performance‑related issues in Linux-based environments.
  • Strong analytical and problem‑solving skills, with the ability to learn unfamiliar technologies quickly and investigate complex technical issues in a structured manner.
  • Good communication and collaboration skills, with the ability to work effectively with engineers and cross‑functional stakeholders across different technical domains and regions.
Preferred Qualification(s)
  • Familiarity with technologies such as BIOS/UEFI, BMC, firmware, PCIe, NVMe, NICs, or hardware telemetry.
  • Proficiency in Python, Go, Bash, or another programming language for production‑grade infrastructure automation, including experience designing, building, or maintaining tools and platforms used in large‑scale production environments.
  • Deep knowledge of Linux administration and troubleshooting, preferably Debian or Ubuntu, combined with strong understanding of server architecture and management technologies such as kernels, drivers, BIOS/UEFI, BMC, Redfish, firmware, PCIe, NVMe, NICs, DPUs, hardware telemetry, and failure diagnostics.
  • Proven hands‑on experience introducing and productionizing large‑scale GPU infrastructure, including ownership of hardware NPI or fleet onboarding across qualification, system integration, deployment, production validation, operational handoff, and post‑launch reliability.
  • Strong understanding of distributed AI workload behavior and performance analysis, including collective communication, multi‑node training, inference serving, GPU scheduling, checkpointing, workload‑related bottlenecks, DCGM, NCCL testing, CUDA profiling, and network‑fabric telemetry.
  • Experience building and operating monitoring, telemetry, hardware management, or automated remediation platforms at substantial scale, with measurable improvements in fleet availability, deployment efficiency, incident reduction, operational efficiency, or reliability.
  • Experience working directly with OEMs, ODMs, component suppliers, or GPU platform vendors throughout qualification, technical escalation, root‑cause analysis, and corrective‑action processes.
  • Experience with one or more advanced infrastructure technologies or engineering areas, such as containerisation and orchestration (e.g., Docker, Kubernetes), infrastructure automation frameworks (e.g., Ansible), AI‑powered automation, AI agents, Large Language Models, Retrieval‑Augmented Generation (RAG), open‑source infrastructure projects, technical publications, patents, or relevant industry standards.
About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an \"Always Day 1\" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Production System Engineer London Regular
Production System Engineer London Regular

ByteDance • Greater London

On-site
GBP 70,000 - 90,000
Diverse and inclusive workplace
Innovative projects
Opportunity for growth
System Software Architect - OS and Kernel Direction
System Software Architect - OS and Kernel Direction

Bytedance • Greater London

On-site
GBP 110,000 - 150,000
Solutions Architect, Byteplus
Solutions Architect, Byteplus

ByteDance • Greater London

On-site
GBP 70,000 - 90,000
Forward Deployed Engineer - BytePlus
Forward Deployed Engineer - BytePlus

Bytedance • Greater London

On-site
GBP 90,000 - 150,000
Solutions Architect, Byteplus
Solutions Architect, Byteplus

ByteDance • Greater London

On-site
GBP 70,000 - 100,000
Solutions Architect, Global Partnerships - Byteplus
Solutions Architect, Global Partnerships - Byteplus

ByteDance • Greater London

On-site
GBP 70,000 - 90,000
AI Solution Architect, BytePlus
AI Solution Architect, BytePlus

ByteDance • Greater London

On-site
GBP 110,000 - 140,000
Data Center Technical Project Manager, Data Center Development
Data Center Technical Project Manager, Data Center Development

ByteDance • Greater London

On-site
GBP 70,000 - 95,000
Production Systems Engineer: Linux, Automation & AI Infra
Production Systems Engineer: Linux, Automation & AI Infra

ByteDance • Greater London

On-site
GBP 70,000 - 110,000
Strategic Account Executive - Innovative Technical Sales (AI)
Strategic Account Executive - Innovative Technical Sales (AI)

ByteDance • Greater London

On-site
GBP 70,000 - 90,000
Diversity and inclusion initiatives
Career growth opportunities
Collaborative work environment