Machine Leaning Performance Engineer (Inference)

Socket.dev

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid working opportunities
Free breakfast, lunch, and snacks
Wellness expense reimbursement
Volunteer opportunities
Social events and celebrations
Workshops and continuous learning

Job summary

Tower Research Capital seeks a gifted ML inference optimization engineer to bridge research and production. You will design and accelerate low-latency inference pipelines, aiming for microsecond latency across heterogeneous hardware in a high-performance trading environment.

As part of the Core Engineering team, you will evaluate platforms, optimize memory hierarchies, and collaborate with ML researchers, HPC and datacenter engineers to deploy scalable, reliable inference solutions that power

Qualifications

  • 2+ years of experience in latency-sensitive DL inference optimization.
  • Deep learning frameworks proficiency, with Python/C++ skills.
  • Experience in GPU kernel development and profiling tools (Nsight, etc).
  • Familiarity with mixed-precision compute and efficient memory usage.

Responsibilities

  • Benchmark inference platforms (CPU, GPU, FPGA) to guide deployments.
  • Optimize execution across memory hierarchies to maximize throughput.
  • Collaborate with Infrastructure to address thermal, power, and operational constraints.
  • Develop highly optimized GPU kernels and integrate performance libraries.
  • Implement model reduction techniques for low-latency, real-time workloads.
  • Collaborate with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers.

Skills

PyTorch/JAX
Python
C++
GPU kernel
Performance profiling
Mixed-precision
ML inference optimization

Tools

Triton
TensorRT
ONNX
IREE
CUTLASS
cuBLAS

Job description

Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.

Tower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.

Engineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.

At Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do - combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.

At Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.

Summary

As part of Tower Research's Core Engineering team, you will bridge the gap between quantitative research and high-performance production systems, architecting inference pipelines that operate at the physical limits of hardware. Your objective will be to drive the speed, efficiency, and reliability of our ML inference pipelines to their absolute limits, ensuring our predictive models consistently achieve microsecond-level latency.

Responsibilities
  • Benchmarking & Strategy:
  • Lead the technical evaluation of diverse inference platforms - ranging across CPUs, GPUs, and FPGAs - to guide Tower's infrastructure deployment decisions.
  • System Architecture Optimization:
  • Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing. You will assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle.
  • Infrastructure & Deployment Feasibility:
  • Collaborate with Infrastructure teams to understand thermal, power, and operational constraints of hardware platforms to design inference strategies for our latency-critical trading strategies that fit within those envelopes.
  • GPU Kernel Development:
  • Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon.
  • Model Optimization & Deployment:
  • Implement advanced model reduction techniques (quantization, pruning, distillation) to ensure compact memory footprints and numerical stability. Prioritize optimization for low-latency, event-level inference workloads to meet real-time trading requirements.
  • Cross-Functional Collaboration:
  • Collaborate closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring to fruition target deployments.
Qualifications
  • 2+ years of experience optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain.
  • ML Frameworks: Deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation.
  • Kernel Development & Optimization Tooling: Proven experience in custom GPU kernel development. Deep familiarity with advanced optimization libraries and compilers (e.g., Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) as well as profiling tools (e.g., Nsight Systems, Nsight Compute).
  • GPU Architecture Mastery: Deep expertise in GPU microarchitecture, encompassing SM execution, warp scheduling, and full memory hierarchy optimization (registers to HBM).
  • Cross-Architecture Benchmarking: Proven record of rigorous, data-driven approach to evaluating inference performance across heterogeneous compute architectures.
  • Bonus: Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs.
  • Prior experience in financial trading is not required.

Anticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus.

Tower’s headquarters are in the historic Equitable Building, right in the heart of NYC’s Financial District and our impact is global, with over a dozen offices around the world.

At Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive - without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.

Our benefits include:

  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch, and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
  • Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats, and celebrations throughout the year
  • Workshops and continuous learning opportunities

At Tower, you’ll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work - together.

Tower Research Capital is an equal opportunity employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Leaning Performance Engineer (Inference)
Machine Leaning Performance Engineer (Inference)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences
Machine Learning Research Engineer New
Machine Learning Research Engineer New

Trading Interview • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off
Hybrid working opportunities
Free breakfast, lunch, and snacks
+5
Machine Learning Research Engineer
Machine Learning Research Engineer

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Free meals in office
Software Engineer, Trading Systems (C++)
Software Engineer, Trading Systems (C++)

Tower Research Capital • New York (NY)

On-site
USD 120,000 - 285,000
Generous paid time off
Hybrid working opportunities
Free breakfast, lunch, and snacks
+2
Software Engineer, Development Tools
Software Engineer, Development Tools

Tower Research Capital • New York (NY)

Hybrid
USD 150,000 - 250,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch, and snacks daily
+3
Senior Systems Engineer
Senior Systems Engineer

Tower Research Capital • New York (NY)

Hybrid
USD 150,000 - 250,000
Generous paid time off
Hybrid working opportunities
Free meals daily
+4
GPU Systems Engineer
GPU Systems Engineer

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Wellness programs
+2
GPU Systems Engineer
GPU Systems Engineer

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Quantitative Trader/Researcher - 2027
Quantitative Trader/Researcher - 2027

Tower Research Capital • New York (NY)

Hybrid
USD 150,000 - 250,000
Generous PTO
Hybrid work
Free meals
+5
Software Engineer, Machine Lifecycle
Software Engineer, Machine Lifecycle

Tower Research Capital • New York (NY)

Hybrid
USD 150,000 - 250,000
Paid time off
Hybrid work
Meals provided
+5