Senior GPU Infrastructure & Production Engineer

ByteDance

San Jose (CA)

On-site

USD 122,000 - 272,000

Full time

11 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance seeks an experienced Senior Production Systems Engineer to lead the introduction and productionization of large-scale GPU infrastructure. You will own platform evaluation, system integration, data center readiness, deployment validation, fleet onboarding, monitoring, and incident response across global data centers.

The role requires strong Linux skills, hardware lifecycle management, and automation experience, with cross-functional collaboration across engineering, data centers, and

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 5+ years of experience in production systems, infrastructure engineering, Site Reliability Engineering, DevOps, hardware systems engineering, or large-scale data center operations.
  • Hands-on experience introducing and productionizing large-scale GPU infrastructure on platforms such as NVIDIA GB200/GB300 NVL72, HGX or DGX B200/B300.
  • Deep knowledge of Linux administration, server architecture, and management tech including BIOS/UEFI, BMC, Redfish, firmware, PCIe, NVMe, NICs, DPUs, telemetry, and diagnostics.
  • Experience deploying distributed AI workloads using containerized/orchestrated environments with CUDA, NCCL, NVLink/NVSwitch, RDMA/InfiniBand, or high-performance Ethernet.

Responsibilities

  • Lead GPU platform introduction, qualification, integration, and production rollout.
  • Define launch criteria and readiness plans for server hardware, firmware, drivers, and software stacks.
  • Collaborate across data center, network, storage, and vendor teams to improve fleet availability and lifecycle management.
  • Diagnose and resolve Linux/hardware/firmware issues; develop benchmarks and health checks.
  • Build automation and telemetry for provisioning, monitoring, and remediation; apply AI for incident triage and remediation.
  • Establish engineering standards, procedures, and long-term support models; mentor engineers.
  • Participate in global on-call rotation and lead incident investigations.

Skills

GPU platforms
Linux administration
Automation
Python
Go
Bash
Hardware lifecycle
Site Reliability Engineering
Incident response

Education

Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or related field

Tools

Kubernetes GPU Operator
Slurm
Ansible
Redfish
PCIe/NVMe

Job description

ByteDance seeks an experienced Senior Production Systems Engineer to lead the introduction and productionization of large-scale GPU infrastructure. You will own platform evaluation, system integration, data center readiness, deployment validation, fleet onboarding, monitoring, and incident response across global data centers.

The role requires strong Linux skills, hardware lifecycle management, and automation experience, with cross-functional collaboration across engineering, data centers, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Graduate Production Systems Engineer – Linux, GPU & AI
Graduate Production Systems Engineer – Linux, GPU & AI

ByteDance • San Jose (CA)

On-site
USD 76,000 - 128,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+1
Graduate Production Systems Engineer - Linux, GPU & AI Ops
Graduate Production Systems Engineer - Linux, GPU & AI Ops

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Graduate Production Systems Engineer: Linux & AI Infra
Graduate Production Systems Engineer: Linux & AI Infra

ByteDance • New York (NY)

On-site
USD 120,000 - 180,000
Graduate GPU AI Platform Engineer — System Optimization
Graduate GPU AI Platform Engineer — System Optimization

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
Senior Cloud Infrastructure Engineer – GPU & DPU
Senior Cloud Infrastructure Engineer – GPU & DPU

Lambda • United States

Remote
USD 180,000 - 260,000
Senior GPU Systems Engineer – Large-Scale AI & HPC
Senior GPU Systems Engineer – Large-Scale AI & HPC

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
Linux System Engineer – Graduate
Linux System Engineer – Graduate

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Senior GPU Fleet Automation Engineer - Bare Metal & Linux
Senior GPU Fleet Automation Engineer - Bare Metal & Linux

Lambda • San Jose (CA)

On-site
USD 150,000 - 210,000
Health Insurance
Dental Coverage
Vision Coverage
+4
Senior System Engineer: Linux Kernel & Large-Scale Systems
Senior System Engineer: Linux Kernel & Large-Scale Systems

ByteDance • San Jose (CA)

On-site
USD 156,000 - 388,000
Medical insurance
Dental insurance
Vision insurance
+9