Low-Latency ML Inference Engineer (GPU/FPGA)

Tower Research Capital

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences

Job summary

Tower Research Capital is hiring for a Core Engineering role focused on building and optimizing ML inference pipelines that push microsecond latency limits. Lead evaluation of CPUs, GPUs, and FPGAs, optimize memory hierarchies, and collaborate with cross-functional teams to deploy low-latency trading workloads.

Required are 2+ years in latency-sensitive DL inference, deep knowledge of PyTorch/JAX, Python/C++, and GPU kernel development.

Qualifications

  • 2+ years optimizing deep learning inference in latency-sensitive or high-throughput environments.
  • Deep expertise in lower-level ML framework development with PyTorch/JAX plus strong Python/C++ skills.
  • Experience in custom GPU kernel development and advanced optimization tooling and profiling.
  • Strong knowledge of GPU microarchitecture and memory hierarchy optimization.

Responsibilities

  • Lead evaluation of diverse inference platforms across CPUs, GPUs, and FPGAs.
  • Analyze and optimize execution across memory hierarchies to maximize performance.
  • Collaborate with Infrastructure to design latency-critical inference strategies.
  • Develop highly optimized GPU kernels and integrate performance libraries.
  • Implement model reduction techniques for low-latency, real-time workloads.
  • Collaborate with ML researchers and engineers to deploy target solutions.

Skills

Python
C++
PyTorch/JAX
GPU optimization
Kernel development
Mixed-precision computing
Data-driven evaluation

Tools

Triton
TensorRT
ONNX
IREE
HLS4ML
cuBLAS
CUTLASS

Job description

Tower Research Capital is hiring for a Core Engineering role focused on building and optimizing ML inference pipelines that push microsecond latency limits. Lead evaluation of CPUs, GPUs, and FPGAs, optimize memory hierarchies, and collaborate with cross-functional teams to deploy low-latency trading workloads.

Required are 2+ years in latency-sensitive DL inference, deep knowledge of PyTorch/JAX, Python/C++, and GPU kernel development.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid Low-Latency ML Inference Engineer
Hybrid Low-Latency ML Inference Engineer

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Free breakfast, lunch, and snacks
Wellness expense reimbursement
+3
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
GPU Systems Engineer - Low-Latency HPC & AI Clusters
GPU Systems Engineer - Low-Latency HPC & AI Clusters

Tower Research Capital • New York (NY)

On-site
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
Free breakfast, lunch & snacks
+4
Low-Latency Quant Developer Intern
Low-Latency Quant Developer Intern

Tower Research Capital • New York (NY)

On-site
Housing accommodation
Free meals
Networking events
+2
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior GPU Infra Engineer — Global Low-Latency ML Inference
Senior GPU Infra Engineer — Global Low-Latency ML Inference

davidjoseph-co • San Francisco (CA)

On-site
USD 150,000 - 250,000
Significant equity
Housing & food (SF hacker house)
Visa sponsorship
+4
Machine Leaning Performance Engineer (Inference)
Machine Leaning Performance Engineer (Inference)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences
Machine Leaning Performance Engineer (Inference)
Machine Leaning Performance Engineer (Inference)

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Free breakfast, lunch, and snacks
Wellness expense reimbursement
+3