Lead GPU Systems Engineer - HPC & AI Infrastructure

Socket.dev

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid working opportunities
Generous PTO
Wellness programs
Free meals
Learning & development

Job summary

Tower Research Capital seeks an accomplished engineer to design, deploy, and scale distributed GPU clusters, building out hardware selection, production operations, and monitoring across thousands of nodes.

You will diagnose bottlenecks across compute, storage, and network layers, collaborating with researchers to benchmark workloads and translate results into speedups. Strong Linux, CUDA/C++, and Python skills are essential.

Qualifications

  • 5+ years engineering large-scale Linux systems in HPC or distributed-infrastructure environments.
  • Deep Linux fundamentals: installation, performance tuning, and kernel-level debugging.
  • Hands-on troubleshooting of distributed GPU workloads with a strong GPU performance model.
  • Experience with GPUDirect RDMA and understanding data movement between GPUs and the network.
  • Proficiency in Python for automation; CUDA or C/C++ experience.
  • Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.
  • Clear communication with researchers, engineers, and vendors.

Responsibilities

  • Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.
  • Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.
  • Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.
  • Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.
  • Own infrastructure projects end to end, from scope and design through implementation and long-term support.
  • Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.

Skills

Linux systems
GPU computing
Python
CUDA/C++
RDMA
GPU troubleshooting
Communication

Tools

Salt
Ansible
Puppet
Chef

Job description

Tower Research Capital seeks an accomplished engineer to design, deploy, and scale distributed GPU clusters, building out hardware selection, production operations, and monitoring across thousands of nodes.

You will diagnose bottlenecks across compute, storage, and network layers, collaborating with researchers to benchmark workloads and translate results into speedups. Strong Linux, CUDA/C++, and Python skills are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Engineer - Low-Latency HPC & AI Clusters
GPU Systems Engineer - Low-Latency HPC & AI Clusters

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Senior GPU Systems Engineer: Scale AI Clusters & HPC
Senior GPU Systems Engineer: Scale AI Clusters & HPC

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior GPU Systems Engineer – Large-Scale AI & HPC
Senior GPU Systems Engineer – Large-Scale AI & HPC

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
Senior GPU HPC Cluster Engineer — Equity Eligible
Senior GPU HPC Cluster Engineer — Equity Eligible

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior HPC Architect — Lead GPU Compute & Scale
Senior HPC Architect — Lead GPU Compute & Scale

NVIDIA AI • Illinois

On-site
USD 184,000 - 357,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Lead HPC Cluster Engineer for GPU AI Compute
Lead HPC Cluster Engineer for GPU AI Compute

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Hybrid HPC Systems Architect - GPU Cloud for AI
Hybrid HPC Systems Architect - GPU Cloud for AI

The Consensus • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Cash compensation
Equity compensation
Health, dental and vision coverage
+1