Senior GPU Production Systems Engineer

ByteDance

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking a Senior Production Systems Engineer to lead the introduction and productionization of high-density GPU infrastructure across its global data centers. You will own platform evaluation, system integration, and fleet onboarding, working closely with AI training and inference environments.

The role demands deep expertise in Linux, hardware lifecycle, and automation, with strong program management and cross-team collaboration across international sites.

Qualifications

  • Bachelor’s degree or equivalent in a related field with 5+ years in production systems or large-scale data center ops.
  • Hands-on experience introducing and productionizing large-scale GPU infrastructure on modern platforms.
  • Strong Linux administration, hardware lifecycle, firmware, and driver knowledge.
  • Experience with containerized AI workloads and orchestration tooling.

Responsibilities

  • Lead evaluation, qualification, and production rollout of next-gen GPU platforms and rack-scale AI infra.
  • Define launch criteria, readiness plans for hardware, firmware, OS, drivers, and tooling.
  • Coordinate across data centers, network, power, cooling, and vendors for fleet readiness.
  • Diagnose complex Linux/hardware issues and develop burn-in, benchmarking, and health-check pipelines.
  • Build scalable automation and telemetry for provisioning, monitoring, and remediation.
  • Provide cross-functional leadership and mentor engineers on global infra programs.
  • Participate in global on-call rotations and lead incident investigations.

Skills

GPU infrastructure
Linux administration
Automation
Python
Incident response
Distributed systems

Education

Bachelor's degree in CS/CE/EE or equivalent

Tools

NVIDIA GPUs
BMC/IPMI/Redfish
Containerization (Docker, Kubernetes)

Job description

ByteDance is seeking a Senior Production Systems Engineer to lead the introduction and productionization of high-density GPU infrastructure across its global data centers. You will own platform evaluation, system integration, and fleet onboarding, working closely with AI training and inference environments.

The role demands deep expertise in Linux, hardware lifecycle, and automation, with strong program management and cross-team collaboration across international sites.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Graduate Production Systems Engineer – Linux, GPU & AI
Graduate Production Systems Engineer – Linux, GPU & AI

ByteDance • San Jose (CA)

On-site
USD 76,000 - 128,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+1
Graduate Production Systems Engineer - Linux, GPU & AI Ops
Graduate Production Systems Engineer - Linux, GPU & AI Ops

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Graduate Production Systems Engineer: Linux & AI Infra
Graduate Production Systems Engineer: Linux & AI Infra

ByteDance • New York (NY)

On-site
USD 120,000 - 180,000
Graduate GPU AI Platform Engineer — System Optimization
Graduate GPU AI Platform Engineer — System Optimization

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Senior Production System Engineer (New York City)
Senior Production System Engineer (New York City)

ByteDance • New York (NY)

On-site
USD 180,000 - 240,000
Senior GPU Systems Engineer – Large-Scale AI & HPC
Senior GPU Systems Engineer – Large-Scale AI & HPC

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
Senior GPU Fabric Architect for AI Cloud
Senior GPU Fabric Architect for AI Cloud

Bitdeer • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Lead GPU Systems Engineer - HPC & AI Infrastructure
Lead GPU Systems Engineer - HPC & AI Infrastructure

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Wellness programs
+2
Infrastructure Systems Engineer Intern – AI & Linux Ops
Infrastructure Systems Engineer Intern – AI & Linux Ops

Bytedance • San Jose (CA)

On-site
USD 34,000 - 62,000