GPU Systems Engineer

Socket.dev

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid working opportunities
Generous PTO
Wellness programs
Free meals
Learning & development

Job summary

Tower Research Capital seeks an accomplished engineer to design, deploy, and scale distributed GPU clusters, building out hardware selection, production operations, and monitoring across thousands of nodes.

You will diagnose bottlenecks across compute, storage, and network layers, collaborating with researchers to benchmark workloads and translate results into speedups. Strong Linux, CUDA/C++, and Python skills are essential.

Qualifications

  • 5+ years engineering large-scale Linux systems in HPC or distributed-infrastructure environments.
  • Deep Linux fundamentals: installation, performance tuning, and kernel-level debugging.
  • Hands-on troubleshooting of distributed GPU workloads with a strong GPU performance model.
  • Experience with GPUDirect RDMA and understanding data movement between GPUs and the network.
  • Proficiency in Python for automation; CUDA or C/C++ experience.
  • Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.
  • Clear communication with researchers, engineers, and vendors.

Responsibilities

  • Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.
  • Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.
  • Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.
  • Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.
  • Own infrastructure projects end to end, from scope and design through implementation and long-term support.
  • Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.

Skills

Linux systems
GPU computing
Python
CUDA/C++
RDMA
GPU troubleshooting
Communication

Tools

Salt
Ansible
Puppet
Chef

Job description

Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.

Tower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.

Engineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.

At Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do — combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.

At Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.

Summary: Trading and research at the firm run around the clock and across the globe, and they run on infrastructure this team designs, builds, and operates.

As part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes.

The role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention.

Responsibilities
  • Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.
  • Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.
  • Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.
  • Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.
  • Own infrastructure projects end to end, from scope and design through implementation and long-term support.
  • Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.
Qualifications
  • 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.
  • Deep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it.
  • Hands‑on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance.
  • Working experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not.
  • Solid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on.
  • Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.
  • Comfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer.
  • Clear communication. You will work daily with researchers, engineers, and vendors.
Nice to Have
  • Experience with the rest of the NVIDIA stack, such as NCCL and NVLink.

Anticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus.

Tower’s headquarters are in the historic Equitable Building, right in the heart of NYC’s Financial District and our impact is global, with over a dozen offices around the world.

At Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive – without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.

Benefits
  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities

At Tower, you’ll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work – together.

Tower Research Capital is an equal opportunity employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer
GPU Systems Engineer

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Software Engineer, GPU Fleet
Software Engineer, GPU Fleet

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals
Software Engineer, GPU Fleet
Software Engineer, GPU Fleet

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous paid time off
Wellness reimbursements
+2
Senior Systems Engineer
Senior Systems Engineer

Tower Research Capital • New York (NY)

Hybrid
USD 150,000 - 250,000
Generous paid time off
Hybrid working opportunities
Free meals daily
+4
HPC Operations Engineer
HPC Operations Engineer

Tower Research Capital • New York (NY)

Hybrid
USD 175,000 - 225,000
Generous paid time off policies
Savings plans and other financial well
Hybrid working opportunities
+5
Machine Leaning Performance Engineer (Inference)
Machine Leaning Performance Engineer (Inference)

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Free breakfast, lunch, and snacks
Wellness expense reimbursement
+3
Machine Leaning Performance Engineer (Inference)
Machine Leaning Performance Engineer (Inference)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences
Software Engineer, Development Tools
Software Engineer, Development Tools

Tower Research Capital • New York (NY)

Hybrid
USD 150,000 - 250,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch, and snacks daily
+3
Machine Learning Research Engineer
Machine Learning Research Engineer

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Free meals in office
Software Engineer, Trading Systems (C++)
Software Engineer, Trading Systems (C++)

Tower Research Capital • New York (NY)

On-site
USD 120,000 - 285,000
Generous paid time off
Hybrid working opportunities
Free breakfast, lunch, and snacks
+2