Get more replies from employers
Send a job-specific resume in minutes.
Long Ridge Partners is seeking a Machine Learning Performance Engineer (Inference) to architect ultra-low-latency inference pipelines for high-frequency trading. You will benchmark workloads across CPU, GPU, and FPGA, drive hardware architecture decisions, and deploy optimized kernels and libraries to maximize throughput and reduce latency in production.
The role focuses on end-to-end performance, including memory hierarchies, interconnects, and thermal/power constraints, with cross-functional
High-Frequency Trading Firm
Compensation: $600,000-1.5 million total
A leading high frequency trading firm is hiring a Machine Learning Performance Engineer to sit at the intersection of quantitative research and high-performance production systems. In his role, you'll architect inference pipelines that operate at the physical limits of hardware, driving the speed, efficiency, and reliability of ML inference so predictive models consistently achieve microsecond-level latency.
GPU usage across the firm's trading teams has grown roughly 100x in the past year as deep learning has moved from a supporting signal to the core of how strategies are built. That growth has outpaced the decision-making around it. Strategies get pushed onto GPUs by default, without anyone systematically asking whether GPU is the right target at all. This role owns that question end to end: benchmark the workload across CPU, GPU, and FPGA, decide the architecture on evidence, then optimize and deploy against it.
You will also have the chance to revisit existing models that never reached production, some of which stalled for hardware or deployment reasons, and run them through different environments to determine where they belong.
This is a role with genuine decision-making scope. Rather than optimizing code for whatever hardware happens to be available, you will determine which hardware the workload should run on in the first place, prove it with data, and then build for it. That combination of architectural judgment and hands-on kernel, and the results are measurable in production almost immediately.
You’ll work on inference at microsecond latency, where the constraints are physical rather than theoretical, and where memory hierarchy, interconnect behavior, thermal envelopes, and fleet utilization all shape the answer.
Benefits include generous paid time off, regional savings and financial wellness plans, hybrid working options, free breakfast, lunch, and snacks daily, in-office wellness experiences and reimbursement for select wellness expenses, company-sponsored sports teams and fitness events, volunteer and charitable giving opportunities, regular social events, and ongoing workshops and learning opportunities.
The culture is collaborative and low on hierarchy, smart, driven people, an open-plan workspace, casual dress, and an environment where the best idea wins.