HPC Operations Engineer

Tower Research Capital

New York (NY)

Híbrido

USD 175.000 - 225.000

Jornada completa

14 días+
Generador de candidaturas

Una candidatura hecha a medida para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Generous paid time off policies
Savings plans and other financial well
Hybrid working opportunities
Free breakfast, lunch, and snacks
In-office wellness experiences
Company-sponsored sports teams
Volunteer opportunities
Social events and continuous learning

Descripción de la vacante

Tower Research Capital in New York seeks an operations-focused HPC support engineer to own day-to-day health of the research compute fleet. You will be first-line support for HPC users across scheduling, compute, storage, and access, handling tickets, triage, provisioning, and routine maintenance.

Sitting with the HPC team, you will monitor queues, node status, service availability, and drive issues to resolution.

Formación

  • Bachelor's degree in CS/engineering or equivalent.
  • 2+ years supporting Linux production environments.
  • Solid Linux administration (RHEL/Ubuntu).
  • Strong written communication and ticketing discipline.
  • Experience working directly with users in a technical support or operations role.

Responsabilidades

  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, and escalate to subject-matter experts when needed.
  • Monitor fleet health (queues, node status, storage, service availability) and act before users report issues.
  • Carry out established operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS installation, configuration, validation, handoff into service).
  • Write and maintain runbooks, knowledge-base articles, and user guides for faster issue resolution.
  • Identify recurring issues and propose practical refinements or automation candidates.

Conocimientos

Linux administration
Customer support
Troubleshooting
Written communication

Educación

Bachelor's degree or equivalent

Herramientas

Slurm
HTCondor
LSF

Descripción del empleo

Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.

Tower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.

Engineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.

At Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do - combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.

At Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.

Summary:

This is an operations role, not a platform engineering one. It is about the daily health of the Research compute fleet: you will be the first line of support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access. The work is transactional by nature, with tickets, triage, provisioning, and maintenance done well, every day. The fleet has grown fast, with close to a thousand machines added recently, and this role exists so that infrastructure gets dedicated, high-standard operational care. You will keep an eye on system health, queues, node status, and service availability; work job failures, scheduler errors, and resource constraints as they come in; and drive every issue to resolution or a clean, well-documented escalation. You will sit inside the HPC team, next to the engineers who build and run the platform. That proximity matters: your diagnostics feed their root-cause work, your runbooks capture what the team learns, and the recurring issues you surface become candidates for automation and permanent fixes. For someone who wants to grow into HPC engineering, this is a strong place to start.

Responsibilities:
  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, and *escalate* to subject-matter experts when a problem extends beyond defined ownership.
  • Monitor fleet health (queues, node status, storage, and service availability) and act on what you see before users have to report it.
  • Carry out established operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS installation, configuration, validation, and handoff into service), and handle reinstalls and decommissions as routine work.
  • Write and maintain runbooks, knowledge-base articles, and user guides so the next occurrence of a problem is faster to fix than the first.
  • Spot recurring issues and propose practical refinements, such as better workflows or automation candidates the HPC team can pick up, so the same ticket stops coming back.
Qualifications:
  • A bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • 2+ years supporting Linux-based production environments.
  • Solid Linux administration fundamentals (RHEL-family and/or Ubuntu).
  • Methodical troubleshooting: you work a problem step by step, know what you have ruled out, and recognize when it is time to elevate.
  • Experience working directly with users in a technical support or operations role.
  • Strong written communication: clear tickets, clear runbooks, clear handoffs.
  • The discipline to follow established processes with genuine attention to detail.
Nice to Have:
  • Enough Bash or Python to script away routine operational tasks.
  • Working knowledge of batch schedulers such as Slurm, HTCondor, or LSF.
  • A good grasp of the plumbing behind networked computing: NFS, automounter, LDAP.
  • Hands-on experience provisioning Linux machines: network boot (PXE), unattended installs (kickstart), or configuration management such as Ansible.
  • Prior exposure to HPC or other large-scale compute environments.
  • Familiarity with monitoring and observability stacks such as Prometheus and Grafana.

Anticipated annual base salary range $175,000-$225,000, plus eligible for discretionary bonus.

Tower’s headquarters are in the historic Equitable Building, right in the heart of NYC’s Financial District and our impact is global, with over a dozen offices around the world.

At Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive - without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.

Our benefits include:
  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities

At Tower, you’ll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work - together.

Tower Research Capital is an equal opportunity employer.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Research Platform Engineer New New York
Research Platform Engineer New New York

Tower Research Capital LLC • New York (NY), Northern (KY)

Híbrido
USD 200.000 - 300.000
Generous PTO
Hybrid work
Free meals
+5
Software Engineer, Machine Lifecycle
Software Engineer, Machine Lifecycle

Tower Research Capital • New York (NY)

Presencial
USD 150.000 - 250.000
Paid time off
Hybrid work
Meals provided
+5
Research Platform Engineer Tower Research Capital · New York, United States 16 hours ago
Research Platform Engineer Tower Research Capital · New York, United States 16 hours ago

Tradermath • New York (NY)

Híbrido
USD 200.000 - 300.000
Generous paid time off
Hybrid working opportunities
Free breakfast, lunch, and snacks
+4
GPU Systems Engineer
GPU Systems Engineer

Tower Research Capital • New York (NY)

Presencial
USD 200.000 - 300.000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Trading Operations Associate
Trading Operations Associate

Tower Research Capital • New York (NY)

Híbrido
USD 120.000 - 200.000
Hybrid working
Free meals
Wellness reimbursement
+3
Software Engineer, GPU Fleet
Software Engineer, GPU Fleet

Tower Research Capital • New York (NY)

Presencial
USD 200.000 - 300.000
Generous PTO
Hybrid work
Free meals
Quantitative Trader/Researcher - 2027
Quantitative Trader/Researcher - 2027

Tower Research Capital • New York (NY)

Presencial
USD 150.000 - 250.000
Generous PTO
Hybrid work
Free meals
+5
Machine Learning Performance Engineer, Training
Machine Learning Performance Engineer, Training

Tower Research Capital • New York (NY)

Híbrido
USD 180.000 - 220.000
Generous PTO
Financial wellness plans
Hybrid work opportunities
+5
Trading Operations Engineer Tower Research Capital · Grand Cayman, United States 10 hours ago
Trading Operations Engineer Tower Research Capital · Grand Cayman, United States 10 hours ago

Tradermath • Northern (KY)

Presencial
USD 120.000 - 180.000
Hybrid work opportunities
Generous paid time off
Wellness reimbursement
Machine Learning Research Engineer
Machine Learning Research Engineer

Tower Research Capital • New York (NY)

Presencial
USD 200.000 - 300.000
Hybrid working opportunities
Generous PTO
Free meals in office