GPU AI Production Systems Engineer

ByteDance

Greater London

On-site

GBP 70,000 - 120,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a hands-on Production Systems Engineer to manage Linux-based server infrastructure, hardware lifecycle, and automation in global data centers. You will work on deployment, monitoring, and maintenance of large-scale fleets, including GPUs, while collaborating with multiple engineering and operations teams to enhance reliability and efficiency.

The role welcomes engineers across levels, with ownership expanding from hands-on engineering to leading complex global initiatives.

Qualifications

  • Bachelor's degree or above in a technical field (CS/CE/EE/IT).
  • 2 years of experience in systems engineering or DevOps/SRE roles or equivalent hands-on work.
  • Strong Linux administration and troubleshooting foundation.
  • Programming or scripting in Python, Bash, Go, or similar for tooling/automation.
  • Hands-on experience troubleshooting Linux-based systems, hardware, storage, networking, and performance.
  • Sharp analytical and problem-solving skills; quick to learn new tech.
  • Good communication and collaboration across teams and regions.

Responsibilities

  • Manage deployment, validation, monitoring, maintenance, and lifecycle of large server fleets (CPU/GPU).
  • Develop automation scripts and tools to reduce manual workload.
  • Troubleshoot Linux systems, hardware, storage, networking, and performance issues.
  • Gain experience with AI infrastructure and GPU server platforms; improve tooling and reliability.
  • Analyze metrics and hardware data to identify trends and risks; document processes.

Skills

Linux administration
Python
Go
Bash
SRE/DevOps
Hardware troubleshooting
Communication

Education

Bachelor's degree in CS/CE/EE/IT

Tools

Docker
Kubernetes
Redfish
Firmware
BMC
PCIe/NVMe

Job description

ByteDance is seeking a hands-on Production Systems Engineer to manage Linux-based server infrastructure, hardware lifecycle, and automation in global data centers. You will work on deployment, monitoring, and maintenance of large-scale fleets, including GPUs, while collaborating with multiple engineering and operations teams to enhance reliability and efficiency.

The role welcomes engineers across levels, with ownership expanding from hands-on engineering to leading complex global initiatives.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production System Engineer
Production System Engineer

ByteDance • Greater London

On-site
GBP 70,000 - 120,000
Production System Engineer London Regular
Production System Engineer London Regular

ByteDance • Greater London

On-site
GBP 70,000 - 90,000
Diverse and inclusive workplace
Innovative projects
Opportunity for growth
GPU Infrastructure Engineer — Scale & Reliability for Production
GPU Infrastructure Engineer — Scale & Reliability for Production

Triwill Group • Greater London

Hybrid
GBP 90,000 - 120,000
GPU Infrastructure Engineer for Scalable AI Inference
GPU Infrastructure Engineer for Scalable AI Inference

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 180,000
System Software Architect - OS and Kernel Direction
System Software Architect - OS and Kernel Direction

ByteDance • Greater London

On-site
GBP 110,000 - 150,000
System Software Architect: OS & Kernel Leadership
System Software Architect: OS & Kernel Leadership

ByteDance • Greater London

On-site
GBP 110,000 - 150,000
Embedded AI Deployment Engineer — Rapid Prototyping in Client Environments
Embedded AI Deployment Engineer — Rapid Prototyping in Client Environments

ByteDance • Greater London

On-site
GBP 90,000 - 140,000
GPU Infra Engineer — Scale, Automation & AI Compute
GPU Infra Engineer — Scale, Automation & AI Compute

OpenAI • Greater London

On-site
GBP 120,000 - 190,000
Lead GPU Infrastructure Architect for Scalable AI Clusters
Lead GPU Infrastructure Architect for Scalable AI Clusters

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000
GPU Infrastructure Lead: Scale, Certification, and Automation
GPU Infrastructure Lead: Scale, Certification, and Automation

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 140,000 - 170,000
Full Benefits