Machine Learning Performance Engineer

Long Ridge Partners

New York (NY)

On-site

USD 600,000 - 1,500,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid working options
Free meals (breakfast, lunch)
Wellness reimbursement
Regular social events
Learning opportunities

Job summary

Long Ridge Partners is seeking a Machine Learning Performance Engineer (Inference) to architect ultra-low-latency inference pipelines for high-frequency trading. You will benchmark workloads across CPU, GPU, and FPGA, drive hardware architecture decisions, and deploy optimized kernels and libraries to maximize throughput and reduce latency in production.

The role focuses on end-to-end performance, including memory hierarchies, interconnects, and thermal/power constraints, with cross-functional

Qualifications

  • 2+ years optimizing deep learning inference in latency-sensitive or high-throughput production environments.
  • Strong Python and C++ skills with understanding of mixed-precision computation.
  • Experience in lower-level ML framework development (PyTorch/JAX) and kernel optimization.
  • Experience building GPU kernels and using optimization libraries and profilers.

Responsibilities

  • Architect and optimize inference pipelines for microsecond-latency trading workloads.
  • Benchmark workloads across CPU/GPU/FPGA and guide hardware deployment decisions.
  • Develop highly optimized kernels and integrate performance libraries.
  • Collaborate with ML researchers, HPC/FPGA/datacenter engineers to productionize deployments.

Skills

Python
C++
PyTorch/JAX
mixed-precision computation
kernel development
profiling

Tools

Triton
TensorRT
ONNX
IREE
HLS4ML
cuBLAS
CUTLASS
Nsight Systems
Nsight Compute

Job description

Machine Learning Performance Engineer (Inference)

High-Frequency Trading Firm

Compensation: $600,000-1.5 million total

About the Opportunity

A leading high frequency trading firm is hiring a Machine Learning Performance Engineer to sit at the intersection of quantitative research and high-performance production systems. In his role, you'll architect inference pipelines that operate at the physical limits of hardware, driving the speed, efficiency, and reliability of ML inference so predictive models consistently achieve microsecond-level latency.

GPU usage across the firm's trading teams has grown roughly 100x in the past year as deep learning has moved from a supporting signal to the core of how strategies are built. That growth has outpaced the decision-making around it. Strategies get pushed onto GPUs by default, without anyone systematically asking whether GPU is the right target at all. This role owns that question end to end: benchmark the workload across CPU, GPU, and FPGA, decide the architecture on evidence, then optimize and deploy against it.

You will also have the chance to revisit existing models that never reached production, some of which stalled for hardware or deployment reasons, and run them through different environments to determine where they belong.

What You'll Do
Benchmarking & Strategy
  • Lead the technical evaluation of inference platforms across CPUs, GPUs, and FPGAs to guide infrastructure deployment decisions
  • Benchmark trading workloads across architectures before compute is committed, and identify where performance gains actually come from — code-level or hardware-level
  • Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing
  • Assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle
Infrastructure & Deployment Feasibility
  • Work with Infrastructure teams to understand the thermal, power, and operational constraints of hardware platforms, and design inference strategies for latency-critical trading strategies that fit within those envelopes
  • Consider the interaction between trading workloads, compute requirements, hardware selection, and fleet utilization
  • Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon
  • Implement advanced model reduction techniques — quantization, pruning, distillation — to ensure compact memory footprints and numerical stability
  • Prioritize optimization for low-latency, event-level inference workloads that meet real-time trading requirements
Cross-Functional Collaboration
  • Partner closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring target deployments to production
What We're Looking For
  • 2+ years optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain
  • ML frameworks: deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation
  • Kernel development and tooling: proven experience building custom GPU kernels, with deep familiarity with optimization libraries and compilers (Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) and profiling tools (Nsight Systems, Nsight Compute)
  • GPU architecture: deep expertise in GPU microarchitecture, including SM execution, warp scheduling, and full memory hierarchy optimization from registers to HBM
  • Prior experience in financial trading is not required.
Nice to Have
  • Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs
Why Join?

This is a role with genuine decision-making scope. Rather than optimizing code for whatever hardware happens to be available, you will determine which hardware the workload should run on in the first place, prove it with data, and then build for it. That combination of architectural judgment and hands-on kernel, and the results are measurable in production almost immediately.

You’ll work on inference at microsecond latency, where the constraints are physical rather than theoretical, and where memory hierarchy, interconnect behavior, thermal envelopes, and fleet utilization all shape the answer.

Benefits include generous paid time off, regional savings and financial wellness plans, hybrid working options, free breakfast, lunch, and snacks daily, in-office wellness experiences and reimbursement for select wellness expenses, company-sponsored sports teams and fitness events, volunteer and charitable giving opportunities, regular social events, and ongoing workshops and learning opportunities.

The culture is collaborative and low on hierarchy, smart, driven people, an open-plan workspace, casual dress, and an environment where the best idea wins.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Machine Learning Engineer (Training & Inference Systems)
Machine Learning Engineer (Training & Inference Systems)

Fintal Partners • New York (NY)

On-site
USD 185,000 - 230,000
Machine Learning Performance Engineer - Quant Research & Trading
Machine Learning Performance Engineer - Quant Research & Trading

Acquire Me • United States

On-site
USD 200,000 - 350,000
Senior ML Engineer, Serving & Optimization
Senior ML Engineer, Serving & Optimization

Selby Jennings • New York (NY)

On-site
USD 180,000 - 250,000
Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences
High-Performance ML Inference Engineer (Hybrid)
High-Performance ML Inference Engineer (Hybrid)

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Hybrid working options
Free meals (breakfast, lunch)
Wellness reimbursement
+2
Machine Learning Hardware Engineer
Machine Learning Hardware Engineer

Fintal Partners • New York (NY)

On-site
USD 140,000 - 190,000
Machine Learning Engineer (Training & Inference Systems)
Machine Learning Engineer (Training & Inference Systems)

Jobzhr • New York (NY)

On-site
USD 180,000 - 280,000
Hardware Machine Learning Engineer
Hardware Machine Learning Engineer

IMC Trading • Chicago (IL)

On-site
USD 200,000 - 225,000
Discretionary bonus
Paid leave
Insurance benefits
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000