HPC (High-Performance Computing) System Admin with strong Slurm, Bright Cluster Manager, HPC C[...]

iShift

United States

Remote

USD 57,308,000 - 85,962,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology solution provider is seeking a remote HPC System Administrator for a contract role lasting 6 to 9 months. The candidate will monitor system health, handle job prioritization using Slurm, and troubleshoot hardware issues in HPC clusters. Strong English communication skills and a significant overlap with US Eastern Time are essential. Ideal candidates should have hands-on experience in Bright Cluster Manager and Linux-based environments. Flexible candidates in the European time zone are welcome if they can accommodate scheduled interactions in EST.

Qualifications

  • Proficient in administering and monitoring clusters through CMSH.
  • Solid understanding of scheduler configuration and job prioritization.
  • Strong background in system management and troubleshooting.
  • Ability to identify hardware faults and perform server-level troubleshooting.
  • Familiarity with Dell iDRAC and HPE iLOM.

Responsibilities

  • Monitor system health and address urgent incidents.
  • Handle Slurm queue status and stuck jobs.
  • Review tickets for triage and prioritization.
  • Address escalated issues and plan system improvements.
  • Plan future system improvements and manage escalations.
  • Coordinate time overlaps with US EST for collaboration.

Skills

Bright Cluster Manager
Slurm
Linux Administration
Hardware Diagnostics
BMC/Remote Management

Job description

Job Title: HPC (High-Performance Computing) System Admin

Location: 100% REMOTE

Employment Type: Contract role - 6 to 9 Months Contract

Capacity and Duration

This role is not full-time. We are looking for a candidate with 50% capacity (20 hours per week) for a duration of six months.

Daily Schedule and Expectations

  • We are looking for a consistent daily presence rather than fragmented hours. The ideal breakdown for their 4-hour workday is as follows:
  • Monitoring system health and addressing urgent incidents.
  • Slurm queue status and stuck jobs
  • Hardware alerts and BMC notifications
  • Ticket Review
  • Triage, prioritization, and initial response to the service desk queue.
  • Project & Escalation (Remaining daily time): Addressing escalated issues and planning future system improvements.
  • Time Zone and Communication
  • The candidate must have a significant overlap with US Eastern Time (EST). It is important to note that the team's technical expert is based in EST, and frequent interaction will be required.
  • We are open to candidates in the European time zone, provided they can maintain the necessary EST overlap.

Because this role involves critical system stability and coordination, strong English communication skills are a requirement.

Job Overview:

We are seeking an experienced HPC System Administrator with hands-on expertise in Bright Cluster Manager, Slurm, Linux environments, and HPC command-line operations. This role involves supporting and maintaining existing production HPC clusters, ensuring stable performance, resolving hardware issues, and assisting users to keep computational workflows running efficiently.

Key Responsibilities & Skills:

  • Bright Cluster Manager: Proficient in administering and monitoring clusters through CMSH, managing system images, and maintaining cluster configurations.
  • Slurm: Solid understanding of scheduler configuration, handling job prioritization, creating policy exceptions, and managing reservations.
  • Linux Administration: Strong background in system management, troubleshooting, and providing technical support to users.
  • Hardware Diagnostics: Ability to identify hardware faults, perform basic server-level troubleshooting, and pinpoint failing components.
  • BMC/Remote Management: Familiarity with Dell iDRAC, HPE iLOM, and Supermicro management interfaces.

Thanks & Best Regards

Piyush Sharma

Recruitment

eMail:Psharma@ishift.net | www.ishift.net

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote HPC Systems Admin (20h/Week Contract)
Remote HPC Systems Admin (20h/Week Contract)

iShift • United States

On-site
USD 57,308,000 - 85,962,000
Systems Administration - HPC Cluster
Systems Administration - HPC Cluster

Metasys Technologies • Newton (MA)

On-site
Senior HPC Specialist (3202-1) Denver, CO
Senior HPC Specialist (3202-1) Denver, CO

ESR Healthcare • Denver (CO)

On-site
USD 90,000 - 130,000
HPC Systems Engineer
HPC Systems Engineer

Radix Trading Experienced Job Board • New York (NY), Chicago (IL)

On-site
USD 120,000 - 150,000
Linux System Administrator (only USC and GC)
Linux System Administrator (only USC and GC)

Ampstek • Michigan

On-site
USD 80,000 - 110,000
HPC Systems Engineer — Remote/Hybrid, Slurm/Linux
HPC Systems Engineer — Remote/Hybrid, Slurm/Linux

Strategic Business Systems, Inc (SBS) • Chantilly (VA)

Hybrid
USD 120,000 - 180,000
Flexible work arrangements
HP-UX System Administrator
HP-UX System Administrator

Raas Infotek • Spring (TX)

On-site
USD 80,000 - 110,000
Linux Systems Administrator
Linux Systems Administrator

Veriipro • Miami (FL)

On-site
USD 90,000 - 140,000
HPC System Administrator
HPC System Administrator

Metasys Technologies • Boston (MA)

Remote
HPC Slurm Administrator - Linux Systems (Contract)
HPC Slurm Administrator - Linux Systems (Contract)

Ampstek • Michigan

On-site
USD 80,000 - 110,000