GPU Systems Engineer - Low-Latency HPC & AI Clusters

Tower Research Capital

New York (NY)

On-site

USD 200,000 - 300,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
Wellness reimbursement
Company-sponsored sports teams and gym
Volunteer opportunities
Social events and learning workshops

Job summary

Tower Research Capital is a leading quantitative trading firm that designs and operates high-performance compute infrastructure for systematic trading. This role sits in R&D, focusing on the compute, storage, and automation behind large-scale trading workloads.

You will join engineers responsible for hundreds of petabytes of storage and thousands of GPU-accelerated nodes, shaping architectures for AI clusters and performance profiling. Collaboration with researchers and vendors is daily.

Qualifications

  • 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.
  • Deep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it.
  • Hands-on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance.
  • Working experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not.
  • Solid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on.
  • Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.
  • Comfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer.
  • Clear communication. You will work daily with researchers, engineers, and vendors.

Responsibilities

  • Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.
  • Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.
  • Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.
  • Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.
  • Own infrastructure projects end to end, from scope and design through implementation and long-term support.
  • Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.

Skills

Linux systems engineering
Distributed GPU workloads
GPU performance modeling
Python automation
C/C++ development

Tools

CUDA
GPUDirect RDMA

Job description

Tower Research Capital is a leading quantitative trading firm that designs and operates high-performance compute infrastructure for systematic trading. This role sits in R&D, focusing on the compute, storage, and automation behind large-scale trading workloads.

You will join engineers responsible for hundreds of petabytes of storage and thousands of GPU-accelerated nodes, shaping architectures for AI clusters and performance profiling. Collaboration with researchers and vendors is daily.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead GPU Systems Engineer - HPC & AI Infrastructure
Lead GPU Systems Engineer - HPC & AI Infrastructure

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Wellness programs
+2
GPU Systems Engineer
GPU Systems Engineer

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
GPU Systems Engineer
GPU Systems Engineer

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Wellness programs
+2
Senior GPU Cluster Architect — Ultra-Low-Latency (1k+ Nodes)
Senior GPU Cluster Architect — Ultra-Low-Latency (1k+ Nodes)

NJF Global Holdings Ltd • New York (NY)

On-site
USD 150,000 - 200,000
Low-Latency ML Inference Engineer (GPU/FPGA)
Low-Latency ML Inference Engineer (GPU/FPGA)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences
AI/ML Systems Engineer Intern - HPC & Low-Latency
AI/ML Systems Engineer Intern - HPC & Low-Latency

Jump Trading • New York (NY)

On-site
USD 270,000 - 330,000
C++ Trading Systems Engineer — Low-Latency, Hybrid NYC
C++ Trading Systems Engineer — Low-Latency, Hybrid NYC

Tower Research Capital • New York (NY)

On-site
USD 120,000 - 285,000
Generous paid time off
Hybrid working opportunities
Free breakfast, lunch, and snacks
+2
Senior GPU Systems Engineer: Scale AI Clusters & HPC
Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)
Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)

NJF Global Holdings Ltd • New York (NY)

On-site
USD 150,000 - 200,000
Low-Latency Quant Developer Intern
Low-Latency Quant Developer Intern

Tower Research Capital • New York (NY)

On-site
Housing accommodation
Free meals
Networking events
+2