Night-Shift Systems Admin Lead - HPC & GPU Clusters

5C Group

United States

On-site

USD 115,000 - 145,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

5C Data Centers is seeking a Systems Administrator Lead (Night Shift) to oversee, troubleshoot, and optimize large-scale HPC and AI infrastructure across data center and cloud environments. You will serve as a senior escalation point for GPU clusters, compute nodes, networking, and storage, operating primarily during the 1AM–10:00 AM EST window.

The role emphasizes hardware and software stack management on Linux, NVIDIA drivers, and out-of-band management, with collaboration across Network

Qualifications

  • Bachelor's degree in CS, Engineering, IT or related field.
  • 7+ years of Linux systems administration in HPC/AI infra.
  • Hands-on with enterprise server hardware and NVIDIA GPU infra.
  • Strong understanding of server hardware, firmware, PCIe, NUMA, memory, storage, and networks.
  • Experience with NVIDIA drivers, CUDA, and GPU-management tools (DCGM, Fabric Manager, nvidia-smi).
  • Knowledge of high-performance networking (InfiniBand, RDMA, RoCE) and out-of-band mgmt tools.
  • Proficiency in Python, Bash, or Ansible; observability tools like Prometheus, Grafana, Datadog.
  • Willlingness to work fixed night shift and on-call escalation.

Responsibilities

  • Administer Linux HPC/AI compute environments, including GPU clusters.
  • Diagnose hardware, OS, driver, firmware, networking, and storage failures; root-cause analysis.
  • Senior escalation point for incidents; restore service in production HPC infra.
  • Support heterogeneous environments - bare-metal, virtualized, containerized, cloud.
  • Install and troubleshoot NVIDIA drivers, CUDA components, DCGM, Fabric Manager.
  • Investigate NCCL, NVLink/NVSwitch, PCIe, ECC, and GPU/CPU issues.
  • Coordinate replacements and RMAs for GPUs and related components.
  • Mentor junior staff; document procedures and shift handoffs.
  • Maintain operational documentation and runbooks; participate in on-call rotas.

Skills

Linux systems
HPC/AI infrastructure
NVIDIA GPUs
NVIDIA DCGM
Out-of-band mgmt
Python scripting
Documentation skills
On-call readiness

Education

Bachelor's degree in Computer Science/Engineering/IT

Tools

NVIDIA DCGM
Fabric Manager
nvidia-smi
IPMI/iDRAC

Job description

5C Data Centers is seeking a Systems Administrator Lead (Night Shift) to oversee, troubleshoot, and optimize large-scale HPC and AI infrastructure across data center and cloud environments. You will serve as a senior escalation point for GPU clusters, compute nodes, networking, and storage, operating primarily during the 1AM–10:00 AM EST window.

The role emphasizes hardware and software stack management on Linux, NVIDIA drivers, and out-of-band management, with collaboration across Network

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Admin Lead Remote HPC/AI (2nd Shift)
Senior Systems Admin Lead Remote HPC/AI (2nd Shift)

5C Group • United States

On-site
USD 115,000 - 145,000
Career Growth
Industry Leadership
Entrepreneurial Culture
+1
AI Data Center Ops Lead — GPU/HPC Infrastructure
AI Data Center Ops Lead — GPU/HPC Infrastructure

Nscale • Town of Norway (WI)

On-site
USD 120,000 - 170,000
Senior HPC Systems Engineer: Linux Clusters, AI & GPU
Senior HPC Systems Engineer: Linux Clusters, AI & GPU

United States Digital Space LLC • Starbase (TX)

On-site
USD 120,000 - 190,000
Senior HPC Support Engineer - GPU Cloud Infra
Senior HPC Support Engineer - GPU Cloud Infra

Neura Market • United States

On-site
USD 150,000 - 190,000
Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
+1
Senior GPU HPC Systems Engineer
Senior GPU HPC Systems Engineer

Acceler8 Talent • Fremont (CA)

On-site
USD 135,000 - 165,000
Comprehensive benefits
Senior HPC GPU Cluster Lead for Deep Learning Infra
Senior HPC GPU Cluster Lead for Deep Learning Infra

NVIDIA • Germany (OH)

On-site
USD 58,000 - 102,000
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
System Engineer
System Engineer

Acceler8 Talent • Fremont (CA)

On-site
USD 135,000 - 165,000
Comprehensive benefits
Systems Admin Supervisor & Helpdesk Lead for AI/HPC Infra
Systems Admin Supervisor & Helpdesk Lead for AI/HPC Infra

Peraton • Chantilly (VA)

On-site
USD 112,000 - 179,000
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits