Remote HPC Systems Administrator - GPU & AI Infra

5C Group

Springfield (OH)

Remote

USD 120,000 - 135,000

Full time

11 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Career growth
Industry leadership
Entrepreneurial culture
Competitive pay

Job summary

5C DATA CENTERS seeks an HPC Systems Administrator to support Linux-based HPC and AI infrastructure across data center and cloud environments in North America.

Based in OH, USA, this remote role focuses on GPU clusters, high-performance networking, storage, and firmware, with opportunities to influence reliability and automation. You will collaborate across teams, participate in on-call rotations, and contribute to scalable HPC platforms.

Qualifications

  • Approximately 5 years of experience in Linux systems administration, data center operations, infrastructure support, HPC, cloud infrastructure, or related technical role.
  • Working knowledge of Linux operating systems and command-line administration.
  • Basic understanding of enterprise server hardware and common server components.
  • Experience troubleshooting hardware or operating-system issues using logs and diagnostic tools.
  • Familiarity with networking fundamentals including IP addressing, DNS, interfaces, routing, and basic connectivity troubleshooting.
  • Familiarity with server hardware concepts including CPU, memory, storage, PCIe devices, NICs, power supplies, and firmware.
  • Exposure to GPU infrastructure or an interest in developing expertise with NVIDIA accelerated computing platforms.
  • Familiarity with out-of-band management technologies such as IPMI, Redfish, iDRAC, iLO, or equivalent tools.
  • Basic scripting or automation experience using Bash, Python, PowerShell, Ansible, or similar technologies.
  • Ability to follow technical procedures and troubleshoot problems methodically.
  • Ability to recognize when an issue requires escalation and clearly communicate collected findings.
  • Strong written communication and documentation skills.
  • Ability to work collaboratively across technical teams.
  • Participate in an on-call rotation.

Responsibilities

  • Support and maintain Linux-based HPC and AI computing environments, including large-scale NVIDIA GPU clusters, while troubleshooting hardware, operating system, networking, storage, driver, and firmware issues across bare-metal, virtualized, containerized, and cloud-hosted platforms.
  • Investigate and resolve infrastructure and compute-node issues, document technical findings, maintain operational runbooks and support procedures, and collaborate with senior engineers to escalate complex problems and improve overall system reliability.
  • Assist with the installation, configuration, validation, monitoring, and troubleshooting of NVIDIA GPU infrastructure, including drivers, firmware, system software, and GPU health management tools, while diagnosing performance, communication, and hardware-related issues.
  • Administer and support enterprise Linux environments, including system services, filesystems, networking, storage, authentication, permissions, SSH, DNS, and NTP, while ensuring compliance with established security and operating system standards.
  • Troubleshoot hardware issues across enterprise server platforms, including memory, PCIe devices, GPUs, NICs, storage systems, power supplies, fans, system boards, and cabling, while supporting firmware updates for BIOS, BMC, GPUs, drives, and other infrastructure components.
  • Utilize out-of-band management tools to perform remote administration, system health monitoring, hardware inventory and log collection, and collaborate with Data Center technicians on hardware replacements, break-fix activities, cabling validation, and post-maintenance testing.
  • Assist with troubleshooting Ethernet, InfiniBand, and RoCE connectivity within HPC and AI environments, including support for NVIDIA/Mellanox network adapters and resolution of link, interface, MTU, packet loss, VLAN, RDMA, and fabric-related issues.
  • Collect and analyze network diagnostic data, elevate complex switching and fabric issues to Network Engineering teams, and validate network performance following infrastructure changes, firmware upgrades, cable replacements, and new system deployments.
  • Participate in incident response activities for production infrastructure environments, troubleshoot technical issues, provide status updates, and elevate incidents in accordance with established support procedures and service-level agreements.
  • Document troubleshooting activities, findings, corrective actions, and resolutions while following change-management processes, maintenance procedures, rollback plans, and on-call support requirements.
  • Collaborate with senior Systems Administrators and engineers to troubleshoot HPC infrastructure issues, follow established operational procedures, and participate in technical reviews, troubleshooting sessions, and post-incident analyses.
  • Contribute to runbooks, knowledge-base articles, and support documentation, share technical findings and lessons learned, and continuously develop expertise in Linux, GPU infrastructure, high-performance networking, storage, and automation technologies.
  • Approximately 5 years of experience in Linux systems administration, data center operations, infrastructure support, HPC, cloud infrastructure, or a related technical role.
  • Working knowledge of Linux operating systems and command-line administration.

Skills

Linux systems administration
Data center operations
HPC infrastructure
Networking fundamentals
Scripting (Bash, Python)
Documentation
On-call experience
Team collaboration

Tools

IPMI / Redfish / iDRAC / iLO
NVIDIA GPU health tools

Job description

5C DATA CENTERS seeks an HPC Systems Administrator to support Linux-based HPC and AI infrastructure across data center and cloud environments in North America.

Based in OH, USA, this remote role focuses on GPU clusters, high-performance networking, storage, and firmware, with opportunities to influence reliability and automation. You will collaborate across teams, participate in on-call rotations, and contribute to scalable HPC platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Systems Administrator
HPC Systems Administrator

5C Group • Springfield (OH)

Remote
USD 120,000 - 135,000
Career growth
Industry leadership
Entrepreneurial culture
+1
Data Center Compute Engineer
Data Center Compute Engineer

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive compensation
Equity opportunity
Comprehensive benefits
+2
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
Remote Technical PM: AI/GPU Cluster Deployments
Remote Technical PM: AI/GPU Cluster Deployments

5C Group • Springfield (OH)

Remote
USD 155,000 - 175,000
Career Growth
Industry Leadership
Entrepreneurial Culture
+1
Remote GPU Cluster Architect - AI Infrastructure Leader
Remote GPU Cluster Architect - AI Infrastructure Leader

Jobgether SRL • United States

Remote
USD 184,000 - 318,000
Medical, dental, vision insurance
Remote work reimbursement
RSUs may be available
+3
Remote HPC Infra Engineer — GPU Clusters
Remote HPC Infra Engineer — GPU Clusters

ElevenLabs • Northern (KY)

Hybrid
USD 140,000 - 210,000
GPU Cloud Systems Administrator | PTO & 401k
GPU Cloud Systems Administrator | PTO & 401k

CyberCoders • Houston (TX)

On-site
USD 120,000 - 140,000
Comprehensive benefits
PTO
401k with match
+1
AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud
AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000
HPC/AI Systems Administrator — Remote/Hybrid
HPC/AI Systems Administrator — Remote/Hybrid

NextSilicon • Minneapolis (MN)

Hybrid
USD 120,000 - 150,000
Hybrid GPU Data Center Engineer: Automation & AI Infra
Hybrid GPU Data Center Engineer: Automation & AI Infra

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits