HPC Operations Pro: Fleet Health & User Support

Career Techniques

New York (NY)

Hybrid

USD 175,000 - 225,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Career Techniques is seeking an operations professional to own the daily health of the Research compute fleet. This role is not platform engineering; you will be first-line support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access.

You will monitor queues, node status, and service availability, triage failures, apply fixes, document procedures, and drive issues to resolution or escalation, while proposing automation opportunities to

Qualifications

  • Bachelor's degree in CS/engineering or equivalent experience.

Responsibilities

  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, escalate when needed.
  • Monitor fleet health (queues, node status, storage, service availability) and act on what you see.
  • Carry out operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS install, configuration, validation, handoff), and handle reinstalls and decommissions.
  • Write and maintain runbooks, knowledge-base articles, and user guides for faster future fixes.
  • Spot recurring issues and propose improvements or automation opportunities for the HPC team to reduce ticket repeats.

Skills

Linux administration
Troubleshooting
Technical support
Documentation
Scripting (Bash/Python)
Team communication

Education

Bachelor's degree or equivalent

Tools

Slurm
HTCondor
LSF
Ansible
PXE
Kickstart
NFS
LDAP
Prometheus
Grafana

Job description

Career Techniques is seeking an operations professional to own the daily health of the Research compute fleet. This role is not platform engineering; you will be first-line support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access.

You will monitor queues, node status, and service availability, triage failures, apply fixes, document procedures, and drive issues to resolution or escalation, while proposing automation opportunities to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Operations Specialist
HPC Operations Specialist

Tower Research Capital • New York (NY)

Hybrid
USD 175,000 - 225,000
Generous paid time off policies
Savings plans and other financial well
Hybrid working opportunities
+5
HPC Operations Engineer
HPC Operations Engineer

Career Techniques • New York (NY)

Hybrid
USD 175,000 - 225,000
Fleet Reliability Engineer for HPC GPU Clusters & Live Ops
Fleet Reliability Engineer for HPC GPU Clusters & Live Ops

CoreWeave • Washington

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+7
HPC Operations Engineer — Power Next-Gen Compute & Equity
HPC Operations Engineer — Power Next-Gen Compute & Equity

NVIDIA • Westford (MA)

On-site
USD 124,000 - 242,000
Fleet Automation Engineer — HPC Infrastructure
Fleet Automation Engineer — HPC Infrastructure

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 110,000 - 160,000
HPC System Administrator for Research Cyberinfrastructure
HPC System Administrator for Research Cyberinfrastructure

Jobtailor • Knoxville (TN)

On-site
USD 75,000 - 110,000
Data Center Operations Lead - HPC & Uptime
Data Center Operations Lead - HPC & Uptime

Core Scientific, Inc • Pecos (TX)

On-site
USD 65,000 - 90,000
HPC Fleet Reliability Engineer
HPC Fleet Reliability Engineer

CoreWeave • Livingston (NJ)

On-site
USD 83,000 - 110,000
Medical insurance
Life Insurance
Disability insurance
+7
Fleet Automation Engineer – HPC Infrastructure
Fleet Automation Engineer – HPC Infrastructure

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lunch stipend
Medical benefits (HDHP)
401(k) match
Systems Admin Supervisor & Helpdesk Lead for AI/HPC Infra
Systems Admin Supervisor & Helpdesk Lead for AI/HPC Infra

Peraton • Chantilly (VA)

On-site
USD 112,000 - 179,000