HPC Operations Engineer

Career Techniques

New York (NY)

Hybrid

USD 175,000 - 225,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Career Techniques is seeking an operations professional to own the daily health of the Research compute fleet. This role is not platform engineering; you will be first-line support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access.

You will monitor queues, node status, and service availability, triage failures, apply fixes, document procedures, and drive issues to resolution or escalation, while proposing automation opportunities to

Qualifications

  • Bachelor's degree in CS/engineering or equivalent experience.

Responsibilities

  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, escalate when needed.
  • Monitor fleet health (queues, node status, storage, service availability) and act on what you see.
  • Carry out operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS install, configuration, validation, handoff), and handle reinstalls and decommissions.
  • Write and maintain runbooks, knowledge-base articles, and user guides for faster future fixes.
  • Spot recurring issues and propose improvements or automation opportunities for the HPC team to reduce ticket repeats.

Skills

Linux administration
Troubleshooting
Technical support
Documentation
Scripting (Bash/Python)
Team communication

Education

Bachelor's degree or equivalent

Tools

Slurm
HTCondor
LSF
Ansible
PXE
Kickstart
NFS
LDAP
Prometheus
Grafana

Job description

This is an operations role, not a platform engineering one. It is about the daily health of the Research compute fleet: you will be the first line of support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access. The work is transactional by nature, with tickets, triage, provisioning, and maintenance done well, every day. You will keep an eye on system health, queues, node status, and service availability; work job failures, scheduler errors, and resource constraints as they come in; and drive every issue to resolution or a clean, well-documented escalation.

Responsibilities:
  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, and elevate to subject-matter experts when a problem extends beyond defined ownership.
  • Monitor fleet health (queues, node status, storage, and service availability) and act on what you see before users have to report it.
  • Carry out established operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS installation, configuration, validation, and handoff into service), and handle reinstalls and decommissions as routine work.
  • Write and maintain runbooks, knowledge-base articles, and user guides so the next occurrence of a problem is faster to fix than the first.
  • Spot recurring issues and propose practical refinements, such as better workflows or automation candidates the HPC team can pick up, so the same ticket stops coming back.
Qualifications:
  • A bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • 2+ years supporting Linux-based production environments.
  • Solid Linux administration fundamentals (RHEL-family and/or Ubuntu).
  • Methodical troubleshooting: you work a problem step by step, know what you have ruled out, and recognize when it is time to elevate.
  • Experience working directly with users in a technical support or operations role.
  • Strong written communication: clear tickets, clear runbooks, clear handoffs.
  • The discipline to follow established processes with genuine attention to detail.
Nice to Have:
  • Enough Bash or python to script away routine operational tasks.
  • Working knowledge of batch schedulers such as Slurm, HTCondor, or LSF.
  • A good grasp of the plumbing behind networked computing: NFS, automounter, LDAP.
  • Hands-on experience provisioning Linux machines: network boot (PXE), unattended installs (kickstart), or configuration management such as Ansible.
  • Prior exposure to HPC or other large-scale compute environments.
  • Familiarity with monitoring and observability stacks such as Prometheus and Grafana.

Comp: $175-225K + Bonus

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

Johns Hopkins University • Baltimore (MD)

On-site
USD 85,500 - 149,800
HPC Systems Engineer
HPC Systems Engineer

Radix Trading Experienced Job Board • Chicago (IL), New York (NY)

On-site
USD 120,000 - 150,000
Technical Expert/Functional Expert (HPC)
Technical Expert/Functional Expert (HPC)

Reflexive Concepts, LLC • Georgetown (MD)

On-site
USD 140,000 - 210,000
HPC Cloud Systems Administrator
HPC Cloud Systems Administrator

RedLine Performance Solutions • College Park (MD)

Remote
USD 120,000 - 150,000
Health benefits
401(k) match
Paid time off
HPC Engineer
HPC Engineer

Tata Consultancy Services • Indianapolis (IN)

On-site
USD 75,000 - 80,000
Discretionary annual incentive
Comprehensive medical coverage (M/D/V,
Family support leaves
+6
Senior HPC Systems Administrator
Senior HPC Systems Administrator

RedLine Performance Solutions • Berkeley (CA)

Remote
USD 140,000 - 190,000
Paid time off
401k match
Health care benefits
HPC Platform Engineer
HPC Platform Engineer

Addison Group • Dallas (TX)

Hybrid
USD 180,000 - 260,000
Medical, dental, and vision insurance
401(k)
25 days PTO
+3
Senior HPC Systems Administrator
Senior HPC Systems Administrator

RedLine • Berkeley (CA)

Remote
USD 140,000 - 190,000
Paid time off
401k match
Health care benefits
HPC Linux Systems Engineer
HPC Linux Systems Engineer

Cadre5 • Knoxville (TN)

On-site
USD 120,000 - 160,000
Excellent medical insurance
Employer-paid benefits
Devops Engineer
Devops Engineer

GIGATEC Engineering • Maryland

On-site
USD 120,000 - 160,000
100% Paid Healthcare
10% 401k fully vested