Hybrid ML Training Performance Engineer — GPU/HPC

Tower Research Capital

New York (NY)

Hybrid

USD 180,000 - 220,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Generous PTO
Financial wellness plans
Hybrid work opportunities
Free meals daily
Wellness reimbursements
Volunteer opportunities
Social events and learning programs
Continuous learning opportunities

Job summary

Tower Research Capital is a leading quantitative trading firm that builds high-performance training systems for machine learning models. In New York, you will bridge research and engineering to accelerate end-to-end training from data ingestion to hardware utilization.

You will work on distributed training optimization, GPU kernel development, and performance benchmarking to enable researchers to iterate across complex models. A hybrid work setup is available in NYC.

Qualifications

  • 3+ years of experience optimizing machine learning training workloads in high-performance, distributed or large-scale computing environments.
  • Deep knowledge of machine learning frameworks such as PyTorch or JAX, including their execution models and distributed-training capabilities.
  • Strong programming skills in Python and C++, with experience developing or optimizing performance-critical systems.
  • Proven experience with GPU kernel development and optimization using CUDA, Triton, CUTLASS, cuBLAS, cuDNN or related libraries.
  • Strong understanding of GPU architecture and memory hierarchy.
  • Experience with distributed-training technologies and communication libraries such as NCCL, DeepSpeed, Megatron-LM, XLA or equivalents.
  • Proficiency with performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler or comparable tools.
  • Understanding of high-performance networking, storage, and accelerator interconnects (InfiniBand, RDMA, NVLink).
  • Demonstrated ability to benchmark heterogeneous compute platforms and make data-driven performance/cost recommendations.

Responsibilities

  • Train and optimize scalable ML models across CPUs/GPUs and accelerator platforms.
  • Design and optimize distributed training strategies (data, tensor, pipeline, model parallelism).
  • Improve communication efficiency across multi-GPU and multi-node environments.
  • Analyze and optimize the full training pipeline from data loading to checkpointing.
  • Develop GPU kernels and framework components for ML workloads.
  • Collaborate with ML researchers, HPC engineers, and systems teams to deliver efficient training systems.

Skills

ML training optimization
PyTorch/JAX
Python
C++
GPU kernel development
NCCL/DeepSpeed
Profiling tools
Distributed training

Tools

CUDA
Triton
CUTLASS
cuBLAS
cuDNN
Nsight
Kubernetes
Slurm
Ray
Megatron-LM

Job description

Tower Research Capital is a leading quantitative trading firm that builds high-performance training systems for machine learning models. In New York, you will bridge research and engineering to accelerate end-to-end training from data ingestion to hardware utilization.

You will work on distributed training optimization, GPU kernel development, and performance benchmarking to enable researchers to iterate across complex models. A hybrid work setup is available in NYC.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Fleet Automation Engineer (Hybrid)
GPU Fleet Automation Engineer (Hybrid)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals
GPU Systems Engineer - Low-Latency HPC & AI Clusters
GPU Systems Engineer - Low-Latency HPC & AI Clusters

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Machine Learning Performance Engineer, Training
Machine Learning Performance Engineer, Training

Tower Research Capital • New York (NY)

Hybrid
USD 180,000 - 220,000
Generous PTO
Financial wellness plans
Hybrid work opportunities
+5
Senior Research Platform Engineer: ML & HPC
Senior Research Platform Engineer: ML & HPC

Tower Research Capital LLC • New York (NY), Northern (KY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals
+5
Performance Engineer - ML Training & CUDA Kernels
Performance Engineer - ML Training & CUDA Kernels

Cohere • New York (NY)

Hybrid
USD 150,000 - 190,000
Lunch stipend
Health benefits
RRSP matching
+5
Senior Research Platform Engineer — HPC & ML
Senior Research Platform Engineer — HPC & ML

Tradermath • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off
Hybrid working opportunities
Free breakfast, lunch, and snacks
+4
GPU-Accelerated ML Training Performance Engineer
GPU-Accelerated ML Training Performance Engineer

Blackrock Neurotech • Salt Lake City (UT)

On-site
USD 120,000 - 190,000
Performance Engineer — GPU & System Optimization
Performance Engineer — GPU & System Optimization

PDT Partners • New York (NY)

Hybrid
USD 90,000 - 130,000
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Senior LLM Training Performance Engineer (Hybrid)
Senior LLM Training Performance Engineer (Hybrid)

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits package
Hybrid work environment