High-Performance ML Inference Engineer (Hybrid)

Long Ridge Partners

New York (NY)

On-site

USD 600,000 - 1,500,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid working options
Free meals (breakfast, lunch)
Wellness reimbursement
Regular social events
Learning opportunities

Job summary

Long Ridge Partners is seeking a Machine Learning Performance Engineer (Inference) to architect ultra-low-latency inference pipelines for high-frequency trading. You will benchmark workloads across CPU, GPU, and FPGA, drive hardware architecture decisions, and deploy optimized kernels and libraries to maximize throughput and reduce latency in production.

The role focuses on end-to-end performance, including memory hierarchies, interconnects, and thermal/power constraints, with cross-functional

Qualifications

  • 2+ years optimizing deep learning inference in latency-sensitive or high-throughput production environments.
  • Strong Python and C++ skills with understanding of mixed-precision computation.
  • Experience in lower-level ML framework development (PyTorch/JAX) and kernel optimization.
  • Experience building GPU kernels and using optimization libraries and profilers.

Responsibilities

  • Architect and optimize inference pipelines for microsecond-latency trading workloads.
  • Benchmark workloads across CPU/GPU/FPGA and guide hardware deployment decisions.
  • Develop highly optimized kernels and integrate performance libraries.
  • Collaborate with ML researchers, HPC/FPGA/datacenter engineers to productionize deployments.

Skills

Python
C++
PyTorch/JAX
mixed-precision computation
kernel development
profiling

Tools

Triton
TensorRT
ONNX
IREE
HLS4ML
cuBLAS
CUTLASS
Nsight Systems
Nsight Compute

Job description

Long Ridge Partners is seeking a Machine Learning Performance Engineer (Inference) to architect ultra-low-latency inference pipelines for high-frequency trading. You will benchmark workloads across CPU, GPU, and FPGA, drive hardware architecture decisions, and deploy optimized kernels and libraries to maximize throughput and reduce latency in production.

The role focuses on end-to-end performance, including memory hierarchies, interconnects, and thermal/power constraints, with cross-functional

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Performance Engineer
Machine Learning Performance Engineer

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Hybrid working options
Free meals (breakfast, lunch)
Wellness reimbursement
+2
Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Low-Latency ML Inference Engineer (GPU/FPGA)
Low-Latency ML Inference Engineer (GPU/FPGA)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences
Hybrid AI Inference Engineer — Kernel & Performance
Hybrid AI Inference Engineer — Kernel & Performance

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
High-Performance ML Inference Engineer
High-Performance ML Inference Engineer

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
ML Hardware Engineer: Custom Hardware Inference
ML Hardware Engineer: Custom Hardware Inference

Fintal Partners • New York (NY)

On-site
USD 140,000 - 190,000
Senior ML Performance Engineer — Ultra-Fast Inference + Equity
Senior ML Performance Engineer — Ultra-Fast Inference + Equity

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
High-Performance Inference Engineer (ML Systems)
High-Performance Inference Engineer (ML Systems)

Garuda Ventures • Hermosa Beach (CA)

On-site
USD 100,000 - 140,000