Machine Learning Performance Engineer, Training

Tower Research Capital

New York (NY)

Hybrid

USD 180,000 - 220,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Generous PTO
Financial wellness plans
Hybrid work opportunities
Free meals daily
Wellness reimbursements
Volunteer opportunities
Social events and learning programs
Continuous learning opportunities

Job summary

Tower Research Capital is a leading quantitative trading firm that builds high-performance training systems for machine learning models. In New York, you will bridge research and engineering to accelerate end-to-end training from data ingestion to hardware utilization.

You will work on distributed training optimization, GPU kernel development, and performance benchmarking to enable researchers to iterate across complex models. A hybrid work setup is available in NYC.

Qualifications

  • 3+ years of experience optimizing machine learning training workloads in high-performance, distributed or large-scale computing environments.
  • Deep knowledge of machine learning frameworks such as PyTorch or JAX, including their execution models and distributed-training capabilities.
  • Strong programming skills in Python and C++, with experience developing or optimizing performance-critical systems.
  • Proven experience with GPU kernel development and optimization using CUDA, Triton, CUTLASS, cuBLAS, cuDNN or related libraries.
  • Strong understanding of GPU architecture and memory hierarchy.
  • Experience with distributed-training technologies and communication libraries such as NCCL, DeepSpeed, Megatron-LM, XLA or equivalents.
  • Proficiency with performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler or comparable tools.
  • Understanding of high-performance networking, storage, and accelerator interconnects (InfiniBand, RDMA, NVLink).
  • Demonstrated ability to benchmark heterogeneous compute platforms and make data-driven performance/cost recommendations.

Responsibilities

  • Train and optimize scalable ML models across CPUs/GPUs and accelerator platforms.
  • Design and optimize distributed training strategies (data, tensor, pipeline, model parallelism).
  • Improve communication efficiency across multi-GPU and multi-node environments.
  • Analyze and optimize the full training pipeline from data loading to checkpointing.
  • Develop GPU kernels and framework components for ML workloads.
  • Collaborate with ML researchers, HPC engineers, and systems teams to deliver efficient training systems.

Skills

ML training optimization
PyTorch/JAX
Python
C++
GPU kernel development
NCCL/DeepSpeed
Profiling tools
Distributed training

Tools

CUDA
Triton
CUTLASS
cuBLAS
cuDNN
Nsight
Kubernetes
Slurm
Ray
Megatron-LM

Job description

Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.

Tower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.

Engineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.

At Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do — combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.

At Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.

Summary

You will bridge the gap between quantitative research and high-performance computing, building and optimizing the systems used to train machine learning models at scale. You will focus on accelerating the end-to-end training lifecycle—from data ingestion and distributed execution to kernel performance and hardware utilization—enabling researchers to iterate more quickly across increasingly complex models and datasets.

Responsibilities
  • Training Performance and Benchmarking
    • Benchmark model-training workloads across CPUs, GPUs, and other accelerator platforms to identify bottlenecks and guide Tower’s compute infrastructure decisions.
    • Develop performance models and standardized benchmarks for measuring throughput, utilization, scalability, and time to convergence.
  • Distributed Training Optimization
    • Design and optimize distributed training strategies, including data, tensor, pipeline, and model parallelism.
    • Improve communication efficiency across multi-GPU and multi-node environments by optimizing collective operations, topology awareness, and computation–communication overlap.
  • End-to-End Training Efficiency
    • Analyze and improve the full training pipeline, including data loading, preprocessing, memory management, forward and backward passes, optimizer execution, checkpointing, and experiment recovery.
    • Identify bottlenecks across compute, memory, storage, networking, and interconnects to increase accelerator utilization and researcher productivity.
  • GPU Kernel and Framework Development
    • Develop and optimize GPU kernels and performance-critical framework components for quantitative machine learning workloads.
    • Integrate specialized libraries, compilers, and execution techniques to improve throughput, memory efficiency, and numerical performance.
  • Model and Numerical Optimization
    • Apply techniques such as mixed-precision training, gradient accumulation, activation checkpointing, operator fusion, and memory-efficient optimizers.
    • Evaluate tradeoffs among training speed, numerical stability, reproducibility, model quality, and infrastructure cost.
  • Training Infrastructure
    • Partner with HPC and infrastructure teams to optimize workload scheduling, resource allocation, observability, fault tolerance, and reproducibility across shared compute environments.
    • Help define the architecture and tooling required to support large-scale experimentation across on-premises and cloud-based infrastructure.
  • Cross-Functional Collaboration
    • Work closely with ML Researchers, Quantitative Researchers, HPC Engineers, Systems Engineers, and hardware specialists to translate research requirements into highly efficient training systems.
Qualifications
  • 3+ years of experience optimizing machine learning training workloads in high-performance, distributed, or large-scale computing environments.
  • Deep knowledge of machine learning frameworks such as PyTorch or JAX, including their execution models, compilation paths, autograd systems, and distributed-training capabilities.
  • Strong programming skills in Python and C++, with experience developing or optimizing performance-critical systems.
  • Proven experience with GPU kernel development and optimization using technologies such as CUDA, Triton, CUTLASS, cuBLAS, cuDNN, or related libraries.
  • Strong understanding of GPU architecture, including streaming multiprocessor execution, warp scheduling, tensor cores, and the memory hierarchy from registers through HBM.
  • Experience with distributed-training technologies and communication libraries such as NCCL, FSDP, DeepSpeed, Megatron-LM, XLA, or equivalent systems.
  • Proficiency with performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, or comparable tracing and profiling platforms.
  • Understanding of high-performance networking, storage, and accelerator interconnects, including technologies such as InfiniBand, RDMA, NVLink, or NVSwitch.
  • Demonstrated ability to benchmark heterogeneous compute platforms and make rigorous, data-driven recommendations about performance, scalability, and cost.
Preferred Qualifications
  • Experience optimizing training workloads for transformer-based, time-series, reinforcement-learning, or other computationally intensive models.
  • Experience with cluster orchestration and scheduling technologies such as Kubernetes, Slurm, Ray, or similar platforms.
  • Familiarity with fault-tolerant distributed training, large-scale checkpointing, experiment reproducibility, and GPU-cluster observability.
  • Practical experience with specialized accelerators, custom hardware, or compiler technologies for machine learning.
  • Prior experience in financial trading is not required.

Anticipated New York annual base salary of $200,000, plus eligible for discretionary bonus.

Benefits

Tower’s headquarters are in the historic Equitable Building, right in the heart of NYC’s Financial District and our impact is global, with over a dozen offices around the world.

At Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive – without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.

  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats and celebrations throughout the year
  • Workshops and continuous learning opportunities

At Tower, you’ll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work – together.

Tower Research Capital is an equal opportunity employer.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Research Engineer
Machine Learning Research Engineer

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Free meals in office
Machine Learning Research Engineer New
Machine Learning Research Engineer New

Trading Interview • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off
Hybrid working opportunities
Free breakfast, lunch, and snacks
+5
Research Platform Engineer Tower Research Capital · New York, United States 16 hours ago
Research Platform Engineer Tower Research Capital · New York, United States 16 hours ago

Tradermath • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off
Hybrid working opportunities
Free breakfast, lunch, and snacks
+4
GPU Systems Engineer
GPU Systems Engineer

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Quantitative Trader/Researcher - 2027
Quantitative Trader/Researcher - 2027

Tower Research Capital • New York (NY)

On-site
USD 150,000 - 250,000
Generous PTO
Hybrid work
Free meals
+5
Software Engineer, Machine Lifecycle
Software Engineer, Machine Lifecycle

Tower Research Capital • New York (NY)

On-site
USD 150,000 - 250,000
Paid time off
Hybrid work
Meals provided
+5
HPC Operations Engineer
HPC Operations Engineer

Tower Research Capital • New York (NY)

On-site
USD 175,000 - 225,000
Generous paid time off policies
Savings plans and other financial well
Hybrid working opportunities
+5
Senior Quantitative Researcher
Senior Quantitative Researcher

Tower Research Capital • New York (NY)

On-site
USD 120,000 - 200,000
Generous paid time off
Hybrid working opportunities
Free breakfast, lunch, and snacks
+2
Trading Operations Associate
Trading Operations Associate

Tower Research Capital • New York (NY)

Hybrid
USD 120,000 - 200,000
Hybrid working
Free meals
Wellness reimbursement
+3
Software Engineer, GPU Fleet
Software Engineer, GPU Fleet

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals