HPC Operations Specialist

Tower Research Capital

New York (NY)

Hybrid

USD 175,000 - 225,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Generous paid time off policies
Savings plans and other financial well
Hybrid working opportunities
Free breakfast, lunch, and snacks
In-office wellness experiences
Company-sponsored sports teams
Volunteer opportunities
Social events and continuous learning

Job summary

Tower Research Capital in New York seeks an operations-focused HPC support engineer to own day-to-day health of the research compute fleet. You will be first-line support for HPC users across scheduling, compute, storage, and access, handling tickets, triage, provisioning, and routine maintenance.

Sitting with the HPC team, you will monitor queues, node status, service availability, and drive issues to resolution.

Qualifications

  • Bachelor's degree in CS/engineering or equivalent.
  • 2+ years supporting Linux production environments.
  • Solid Linux administration (RHEL/Ubuntu).
  • Strong written communication and ticketing discipline.
  • Experience working directly with users in a technical support or operations role.

Responsibilities

  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, and escalate to subject-matter experts when needed.
  • Monitor fleet health (queues, node status, storage, service availability) and act before users report issues.
  • Carry out established operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS installation, configuration, validation, handoff into service).
  • Write and maintain runbooks, knowledge-base articles, and user guides for faster issue resolution.
  • Identify recurring issues and propose practical refinements or automation candidates.

Skills

Linux administration
Customer support
Troubleshooting
Written communication

Education

Bachelor's degree or equivalent

Tools

Slurm
HTCondor
LSF

Job description

Tower Research Capital in New York seeks an operations-focused HPC support engineer to own day-to-day health of the research compute fleet. You will be first-line support for HPC users across scheduling, compute, storage, and access, handling tickets, triage, provisioning, and routine maintenance.

Sitting with the HPC team, you will monitor queues, node status, service availability, and drive issues to resolution.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Operations Pro: Fleet Health & User Support
HPC Operations Pro: Fleet Health & User Support

Career Techniques • New York (NY)

Hybrid
USD 175,000 - 225,000
HPC Operations Engineer
HPC Operations Engineer

Tower Research Capital • New York (NY)

Hybrid
USD 175,000 - 225,000
Generous paid time off policies
Savings plans and other financial well
Hybrid working opportunities
+5
HPC Operations Engineer
HPC Operations Engineer

Career Techniques • New York (NY)

Hybrid
USD 175,000 - 225,000
HPC Support Engineer, Research Computing
HPC Support Engineer, Research Computing

Columbia University Irving Medical Center • New York (NY)

On-site
USD 90,000 - 100,000
HPC Infrastructure Lead & Automation Architect
HPC Infrastructure Lead & Automation Architect

P2P • New York (NY)

On-site
USD 125,000 - 150,000
Discretionary bonus eligibility
Medical, dental, and vision insurance
Paid vacation plus paid holidays
+2
HPC Operations Engineer — Power Next-Gen Compute & Equity
HPC Operations Engineer — Power Next-Gen Compute & Equity

NVIDIA • Westford (MA)

On-site
USD 124,000 - 242,000
GPU Systems Engineer - Low-Latency HPC & AI Clusters
GPU Systems Engineer - Low-Latency HPC & AI Clusters

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Support Technician, HPC
Support Technician, HPC

Columbia University Irving Medical Center • New York (NY)

On-site
USD 90,000 - 100,000
HPC Storage Engineer for Large-Scale AI Workloads
HPC Storage Engineer for Large-Scale AI Workloads

Hudson River Trading • Chicago (IL), New York (NY), Austin (TX)

On-site
USD 150,000 - 250,000
HPC Data Center Operations Lead — On‑Site, Travel
HPC Data Center Operations Lead — On‑Site, Travel

P2P • New York (NY)

On-site
USD 120,000 - 160,000