HPC Operations Engineer

Career Techniques

New York (NY)

Hybrid

USD 175,000 - 225,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Career Techniques is seeking an operations professional to own the daily health of the Research compute fleet. This role is not platform engineering; you will be first-line support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access.

You will monitor queues, node status, and service availability, triage failures, apply fixes, document procedures, and drive issues to resolution or escalation, while proposing automation opportunities to

Qualifications

  • Bachelor's degree in CS/engineering or equivalent experience.

Responsibilities

  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, escalate when needed.
  • Monitor fleet health (queues, node status, storage, service availability) and act on what you see.
  • Carry out operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS install, configuration, validation, handoff), and handle reinstalls and decommissions.
  • Write and maintain runbooks, knowledge-base articles, and user guides for faster future fixes.
  • Spot recurring issues and propose improvements or automation opportunities for the HPC team to reduce ticket repeats.

Skills

Linux administration
Troubleshooting
Technical support
Documentation
Scripting (Bash/Python)
Team communication

Education

Bachelor's degree or equivalent

Tools

Slurm
HTCondor
LSF
Ansible
PXE
Kickstart
NFS
LDAP
Prometheus
Grafana

Job description

This is an operations role, not a platform engineering one. It is about the daily health of the Research compute fleet: you will be the first line of support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access. The work is transactional by nature, with tickets, triage, provisioning, and maintenance done well, every day. You will keep an eye on system health, queues, node status, and service availability; work job failures, scheduler errors, and resource constraints as they come in; and drive every issue to resolution or a clean, well-documented escalation.

Responsibilities:
  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, and elevate to subject-matter experts when a problem extends beyond defined ownership.
  • Monitor fleet health (queues, node status, storage, and service availability) and act on what you see before users have to report it.
  • Carry out established operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS installation, configuration, validation, and handoff into service), and handle reinstalls and decommissions as routine work.
  • Write and maintain runbooks, knowledge-base articles, and user guides so the next occurrence of a problem is faster to fix than the first.
  • Spot recurring issues and propose practical refinements, such as better workflows or automation candidates the HPC team can pick up, so the same ticket stops coming back.
Qualifications:
  • A bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • 2+ years supporting Linux-based production environments.
  • Solid Linux administration fundamentals (RHEL-family and/or Ubuntu).
  • Methodical troubleshooting: you work a problem step by step, know what you have ruled out, and recognize when it is time to elevate.
  • Experience working directly with users in a technical support or operations role.
  • Strong written communication: clear tickets, clear runbooks, clear handoffs.
  • The discipline to follow established processes with genuine attention to detail.
Nice to Have:
  • Enough Bash or python to script away routine operational tasks.
  • Working knowledge of batch schedulers such as Slurm, HTCondor, or LSF.
  • A good grasp of the plumbing behind networked computing: NFS, automounter, LDAP.
  • Hands-on experience provisioning Linux machines: network boot (PXE), unattended installs (kickstart), or configuration management such as Ansible.
  • Prior exposure to HPC or other large-scale compute environments.
  • Familiarity with monitoring and observability stacks such as Prometheus and Grafana.

Comp: $175-225K + Bonus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC System Administrator
HPC System Administrator

Cybotic System • Savannah (GA)

On-site
USD 90,000 - 150,000
Senior HPC Systems Engineer
Senior HPC Systems Engineer

Autonomai Recruitment • Chicago (IL)

On-site
USD 110,000 - 170,000
Senior HPC & Infrastructure Engineer
Senior HPC & Infrastructure Engineer

Soni • Cherry Hill Township (NJ)

On-site
USD 135,000 - 155,000
Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

Johns Hopkins University • Baltimore (MD)

On-site
USD 85,000 - 150,000
HPC Systems Engineer
HPC Systems Engineer

Radix Trading Experienced Job Board • New York (NY), Chicago (IL)

On-site
USD 120,000 - 150,000
HPC Operations Pro: Fleet Health & User Support
HPC Operations Pro: Fleet Health & User Support

Career Techniques • New York (NY)

Hybrid
USD 175,000 - 225,000
Mid HPC Engineer
Mid HPC Engineer

Jobtailor • Beavercreek (OH)

On-site
USD 65,000 - 90,000
HPC System Administrator
HPC System Administrator

Jobtailor • Irving (TX)

On-site
USD 85,000 - 115,000
Compute Platform Engineer
Compute Platform Engineer

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 120,000 - 180,000
Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

The Johns Hopkins University • Baltimore (MD)

On-site
USD 85,000 - 150,000