HPC Systems Administrator

5C Group

Springfield (OH)

Remote

USD 120,000 - 135,000

Full time

11 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Career growth
Industry leadership
Entrepreneurial culture
Competitive pay

Job summary

5C DATA CENTERS seeks an HPC Systems Administrator to support Linux-based HPC and AI infrastructure across data center and cloud environments in North America.

Based in OH, USA, this remote role focuses on GPU clusters, high-performance networking, storage, and firmware, with opportunities to influence reliability and automation. You will collaborate across teams, participate in on-call rotations, and contribute to scalable HPC platforms.

Qualifications

  • Approximately 5 years of experience in Linux systems administration, data center operations, infrastructure support, HPC, cloud infrastructure, or related technical role.
  • Working knowledge of Linux operating systems and command-line administration.
  • Basic understanding of enterprise server hardware and common server components.
  • Experience troubleshooting hardware or operating-system issues using logs and diagnostic tools.
  • Familiarity with networking fundamentals including IP addressing, DNS, interfaces, routing, and basic connectivity troubleshooting.
  • Familiarity with server hardware concepts including CPU, memory, storage, PCIe devices, NICs, power supplies, and firmware.
  • Exposure to GPU infrastructure or an interest in developing expertise with NVIDIA accelerated computing platforms.
  • Familiarity with out-of-band management technologies such as IPMI, Redfish, iDRAC, iLO, or equivalent tools.
  • Basic scripting or automation experience using Bash, Python, PowerShell, Ansible, or similar technologies.
  • Ability to follow technical procedures and troubleshoot problems methodically.
  • Ability to recognize when an issue requires escalation and clearly communicate collected findings.
  • Strong written communication and documentation skills.
  • Ability to work collaboratively across technical teams.
  • Participate in an on-call rotation.

Responsibilities

  • Support and maintain Linux-based HPC and AI computing environments, including large-scale NVIDIA GPU clusters, while troubleshooting hardware, operating system, networking, storage, driver, and firmware issues across bare-metal, virtualized, containerized, and cloud-hosted platforms.
  • Investigate and resolve infrastructure and compute-node issues, document technical findings, maintain operational runbooks and support procedures, and collaborate with senior engineers to escalate complex problems and improve overall system reliability.
  • Assist with the installation, configuration, validation, monitoring, and troubleshooting of NVIDIA GPU infrastructure, including drivers, firmware, system software, and GPU health management tools, while diagnosing performance, communication, and hardware-related issues.
  • Administer and support enterprise Linux environments, including system services, filesystems, networking, storage, authentication, permissions, SSH, DNS, and NTP, while ensuring compliance with established security and operating system standards.
  • Troubleshoot hardware issues across enterprise server platforms, including memory, PCIe devices, GPUs, NICs, storage systems, power supplies, fans, system boards, and cabling, while supporting firmware updates for BIOS, BMC, GPUs, drives, and other infrastructure components.
  • Utilize out-of-band management tools to perform remote administration, system health monitoring, hardware inventory and log collection, and collaborate with Data Center technicians on hardware replacements, break-fix activities, cabling validation, and post-maintenance testing.
  • Assist with troubleshooting Ethernet, InfiniBand, and RoCE connectivity within HPC and AI environments, including support for NVIDIA/Mellanox network adapters and resolution of link, interface, MTU, packet loss, VLAN, RDMA, and fabric-related issues.
  • Collect and analyze network diagnostic data, elevate complex switching and fabric issues to Network Engineering teams, and validate network performance following infrastructure changes, firmware upgrades, cable replacements, and new system deployments.
  • Participate in incident response activities for production infrastructure environments, troubleshoot technical issues, provide status updates, and elevate incidents in accordance with established support procedures and service-level agreements.
  • Document troubleshooting activities, findings, corrective actions, and resolutions while following change-management processes, maintenance procedures, rollback plans, and on-call support requirements.
  • Collaborate with senior Systems Administrators and engineers to troubleshoot HPC infrastructure issues, follow established operational procedures, and participate in technical reviews, troubleshooting sessions, and post-incident analyses.
  • Contribute to runbooks, knowledge-base articles, and support documentation, share technical findings and lessons learned, and continuously develop expertise in Linux, GPU infrastructure, high-performance networking, storage, and automation technologies.
  • Approximately 5 years of experience in Linux systems administration, data center operations, infrastructure support, HPC, cloud infrastructure, or a related technical role.
  • Working knowledge of Linux operating systems and command-line administration.

Skills

Linux systems administration
Data center operations
HPC infrastructure
Networking fundamentals
Scripting (Bash, Python)
Documentation
On-call experience
Team collaboration

Tools

IPMI / Redfish / iDRAC / iLO
NVIDIA GPU health tools

Job description

5C DATA CENTERS

Join the Future of Digital Infrastructure

Are you a passionate HPC Systems Administrator looking to make a meaningful impact? We're building the next generation of digital infrastructure powering hyperscalers, AI innovation, and high-performance computing across North America.

JOB TITLE HPC Systems Administrator

DEPARTMENT Cloud Operations

LOCATION OH, USA

WORK ARRANGEMENT Remote

SALARY RANGE $120,000 – $135,000

Role Summary

As an HPC Systems Administrator, you will support the day-to-day operation, maintenance, and troubleshooting of HPC and AI infrastructure across data center and cloud environments.

This role focuses primarily on the infrastructure and hardware layer, including Linux operating systems, GPU servers, high-speed networking, storage connectivity, firmware, drivers, and out-of-band management.

You will work alongside senior Systems Administrators, Network Engineering, Data Center Operations, Deployment Engineering, vendors, and other technical teams to troubleshoot infrastructure issues, restore systems to service, perform routine maintenance, and improve the reliability of production HPC environments.

This position is well suited for someone with a foundation in Linux systems administration or data center infrastructure who wants to develop deeper expertise in HPC, GPU computing, high-performance networking, and large-scale AI infrastructure.

How We Work at 5C

Our core values guide how we collaborate, make decisions, support one another, and serve our customers. We're looking for people who embrace them and help us build something great.

What You Will Do
HPC Infrastructure Operations
  • Support and maintain Linux-based HPC and AI computing environments, including large-scale NVIDIA GPU clusters, while troubleshooting hardware, operating system, networking, storage, driver, and firmware issues across bare-metal, virtualized, containerized, and cloud-hosted platforms.
  • Investigate and resolve infrastructure and compute-node issues, document technical findings, maintain operational runbooks and support procedures, and collaborate with senior engineers to **escalate** complex problems and improve overall system reliability.
GPU and Accelerated Computing Systems
  • Assist with the installation, configuration, validation, monitoring, and troubleshooting of NVIDIA GPU infrastructure, including drivers, firmware, system software, and GPU health management tools, while diagnosing performance, communication, and hardware-related issues.
  • Perform post-maintenance system validation, collect and analyze diagnostic data for GPU and infrastructure faults, and coordinate hardware replacement and RMA activities for GPUs, system boards, NICs, power supplies, and other server components.
Linux System Administration
  • Administer and support enterprise Linux environments, including system services, filesystems, networking, storage, authentication, permissions, SSH, DNS, and NTP, while ensuring compliance with established security and operating system standards.
  • Troubleshoot operating system, CPU, memory, device-discovery, and kernel-related issues, and collect diagnostic data to investigate system crashes, hardware failures, and performance-related incidents.
Server Hardware, Firmware, and Out-of-Band Management
  • Troubleshoot hardware issues across enterprise server platforms, including memory, PCIe devices, GPUs, NICs, storage systems, power supplies, fans, system boards, and cabling, while supporting firmware updates for BIOS, BMC, GPUs, drives, and other infrastructure components.
  • Utilize out-of-band management tools to perform remote administration, system health monitoring, hardware inventory and log collection, and collaborate with Data Center technicians on hardware replacements, break-fix activities, cabling validation, and post-maintenance testing.
High-Performance Networking
  • Assist with troubleshooting Ethernet, InfiniBand, and RoCE connectivity within HPC and AI environments, including support for NVIDIA/Mellanox network adapters and resolution of link, interface, MTU, packet loss, VLAN, RDMA, and fabric-related issues.
  • Collect and analyze network diagnostic data, elevate complex switching and fabric issues to Network Engineering teams, and validate network performance following infrastructure changes, firmware upgrades, cable replacements, and new system deployments.
Incident, Problem, and Change Management
  • Participate in incident response activities for production infrastructure environments, troubleshoot technical issues, provide status updates, and elevate incidents in accordance with established support procedures and service-level agreements.
  • Document troubleshooting activities, findings, corrective actions, and resolutions while following change-management processes, maintenance procedures, rollback plans, and on-call support requirements.
Collaboration and Development
  • Collaborate with senior Systems Administrators and engineers to troubleshoot HPC infrastructure issues, follow established operational procedures, and participate in technical reviews, troubleshooting sessions, and post-incident analyses.
  • Contribute to runbooks, knowledge-base articles, and support documentation, share technical findings and lessons learned, and continuously develop expertise in Linux, GPU infrastructure, high-performance networking, storage, and automation technologies.
What You Bring
  • Approximately 5 years of experience in Linux systems administration, data center operations, infrastructure support, HPC, cloud infrastructure, or a related technical role. Equivalent hands-on experience, education, or technical training may also be considered.
  • Working knowledge of Linux operating systems and command-line administration.
  • Basic understanding of enterprise server hardware and common server components.
  • Experience troubleshooting hardware or operating-system issues using logs and diagnostic tools.
  • Familiarity with networking fundamentals including IP addressing, DNS, interfaces, routing, and basic connectivity troubleshooting.
  • Familiarity with server hardware concepts including CPU, memory, storage, PCIe devices, NICs, power supplies, and firmware.
  • Exposure to GPU infrastructure or an interest in developing expertise with NVIDIA accelerated computing platforms.
  • Familiarity with out-of-band management technologies such as IPMI, Redfish, iDRAC, iLO, or equivalent tools.
  • Basic scripting or automation experience using Bash, Python, PowerShell, Ansible, or similar technologies.
  • Ability to follow technical procedures and troubleshoot problems methodically.
  • Ability to recognize when an issue requires escalation and clearly communicate collected findings.
  • Strong written communication and documentation skills.
  • Ability to work collaboratively across technical teams.
  • Participate in an on-call rotation.
Why Join 5C Data Centers?

At 5C, we believe great people build great companies. You'll build a rewarding career while helping shape the future of digital infrastructure - one of the fastest-growing industries in the world.

Career Growth Build a rewarding career in one of the world's fastest-growing industries.

Industry Leadership Help power the infrastructure behind AI and high-performance computing.

Entrepreneurial Culture Your ideas matter - we empower employees to help shape our future.

Comprehensive Rewards Competitive pay plus meaningful, lasting impact on the work you do.

Life at 5C

We're more than a workplace - we're a team of builders, innovators, and problem-solvers united by a shared purpose: creating infrastructure that powers the technologies transforming our world. Your voice matters here, and we encourage fresh ideas at every level.

5C Data Centers is an equal opportunity employer.

We celebrate diversity and are committed to creating an inclusive environment where everyone can thrive. 5C evaluates qualified applicants without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity or expression, disability status, or any other legally protected characteristic.

#LI-AJ1

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Project Manager, GPU Infrastructure Deployment
Technical Project Manager, GPU Infrastructure Deployment

5C Group • Springfield (OH)

Remote
USD 155,000 - 175,000
Career Growth
Industry Leadership
Entrepreneurial Culture
+1
Network Engineer
Network Engineer

5C Group • Springfield (OH)

On-site
USD 105,000 - 145,000
Sr. Manager, Compute Hardware Sourcing & Operations
Sr. Manager, Compute Hardware Sourcing & Operations

5C • United States

On-site
USD 160,000 - 260,000
Chief Engineer Mechanical & Electrical
Chief Engineer Mechanical & Electrical

5C Group • Phoenix (AZ)

On-site
USD 175,000 - 210,000
Network Engineer
Network Engineer

5C • Springfield (OH)

Remote
USD 105,000 - 145,000
Sr. Manager, Capital Equipment Procurement
Sr. Manager, Capital Equipment Procurement

5C • Springfield (OH)

On-site
USD 180,000 - 240,000
Financial Analyst, FP&A
Financial Analyst, FP&A

5C • Kentucky

On-site
USD 85,000 - 105,000
Remote HPC Systems Administrator - GPU & AI Infra
Remote HPC Systems Administrator - GPU & AI Infra

5C Group • Springfield (OH)

Remote
USD 120,000 - 135,000
Career growth
Industry leadership
Entrepreneurial culture
+1
HPC Linux Systems Engineer
HPC Linux Systems Engineer

Cadre5 • Knoxville (TN)

On-site
USD 120,000 - 160,000
Excellent medical insurance
Employer-paid benefits
Critical Facilities Maintenance Engineer
Critical Facilities Maintenance Engineer

5C • Springfield (OH)

On-site
USD 90,000 - 120,000