Microsecond ML Inference Architect

Long Ridge Partners

New York (NY)

On-site

USD 600,000 - 1,500,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Generous PTO
Hybrid work options
Free breakfast, lunch, snacks

Job summary

Long Ridge Partners is seeking a Machine Learning Performance Engineer (Inference) in New York to architect low-latency ML inference pipelines for high-frequency trading. You will benchmark CPU, GPU, and FPGA workloads, optimize kernels, and push models into real-time production with strict latency requirements.

You will collaborate with ML researchers, HPC, FPGA, and Datacenter teams to ensure efficient, scalable deployments across hardware platforms.

Qualifications

  • 2+ years optimizing deep learning inference in latency-sensitive or high-throughput environments.
  • Expertise in lower-level ML framework development (PyTorch/JAX) with solid Python/C++ skills.
  • Experience building custom GPU kernels and using optimization libraries (TensorRT, Triton) for low-latency workloads.

Responsibilities

  • Lead evaluation of inference platforms (CPU/GPU/FPGA) for latency and throughput.
  • Benchmark workloads across architectures and identify performance gains from code vs hardware.
  • Develop optimized kernels and integrate libraries to maximize throughput on silicon.
  • Design and implement model reduction techniques to meet real-time trading latency targets.
  • Collaborate with ML researchers, HPC, FPGA, and Datacenter engineers to productionize deployments.

Skills

Python programming
C++ programming
Mixed-precision compute
Latency optimization
High-throughput systems

Tools

PyTorch
JAX
Triton
TensorRT
ONNX
IREE
HLS4ML
cuBLAS
CUTLASS
Nsight Systems
Nsight Compute

Job description

Long Ridge Partners is seeking a Machine Learning Performance Engineer (Inference) in New York to architect low-latency ML inference pipelines for high-frequency trading. You will benchmark CPU, GPU, and FPGA workloads, optimize kernels, and push models into real-time production with strict latency requirements.

You will collaborate with ML researchers, HPC, FPGA, and Datacenter teams to ensure efficient, scalable deployments across hardware platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
High-Performance ML Inference Engineer for Markets
High-Performance ML Inference Engineer for Markets

Fintal Partners • New York (NY)

On-site
USD 140,000 - 220,000
Low-Latency ML Inference Engineer (GPU/FPGA)
Low-Latency ML Inference Engineer (GPU/FPGA)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences
ML Hardware Engineer: Custom Hardware Inference
ML Hardware Engineer: Custom Hardware Inference

Fintal Partners • New York (NY)

On-site
USD 140,000 - 190,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Generous PTO
Hybrid work options
Free breakfast, lunch, snacks
Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Fintal Partners • Chicago (IL)

On-site
USD 150,000 - 230,000
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Machine Learning Engineer - Inference
Machine Learning Engineer - Inference

Fintal Partners • New York (NY)

On-site
USD 140,000 - 220,000
ML Inference Engineer - Low-Latency GPU Systems
ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000