Low-Latency ML Inference Engineer

Career Techniques

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Career Techniques in New York, NY seeks a hands-on ML inference engineer to optimize production-grade models across CPUs, GPUs, and FPGAs. You will lead platform evaluations, tune kernels, and push memory and interconnect performance to meet latency requirements.

Collaborate with ML researchers, HPC and datacenter teams to deploy compact, low-latency inference workloads, using tools like Triton, TensorRT and Nsight, with a focus on reliability and scalable throughput.

Qualifications

  • 2+ years optimizing DL inference in latency-sensitive environments.
  • Deep expertise in PyTorch/JAX with Python/C++.
  • Experience in custom GPU kernel development and libraries for performance.
  • Strong knowledge of GPU architecture and memory hierarchies.
  • Experience benchmarking across heterogeneous compute architectures.

Responsibilities

  • Lead evaluation of inference platforms across CPUs, GPUs, and FPGAs.
  • Analyze and optimize execution across deep memory hierarchies.
  • Collaborate with infrastructure teams on latency-critical strategy.
  • Develop highly optimized GPU kernels and integrate libraries.
  • Implement model compression for low-latency inference.
  • Work with ML researchers, HPC/FPGA/datacenter teams on deployments.

Skills

PyTorch
JAX
Python
C++
Mixed-precision
GPU kernel development
TensorRT
Profiling tools
CuBLAS/CUTLASS
FPGAs/ASICs

Tools

Triton
TensorRT
ONNX
IREE
HLS4ML
cuBLAS
CUTLASS
Nsight Systems
Nsight Compute

Job description

Career Techniques in New York, NY seeks a hands-on ML inference engineer to optimize production-grade models across CPUs, GPUs, and FPGAs. You will lead platform evaluations, tune kernels, and push memory and interconnect performance to meet latency requirements.

Collaborate with ML researchers, HPC and datacenter teams to deploy compact, low-latency inference workloads, using tools like Triton, TensorRT and Nsight, with a focus on reliability and scalable throughput.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Microsecond ML Inference Architect
Microsecond ML Inference Architect

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Generous PTO
Hybrid work options
Free breakfast, lunch, snacks
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
ML Inference Engineer - Low-Latency GPU Systems
ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Scale Low-Latency ML Inference Engineer (GPU/CUDA)
Scale Low-Latency ML Inference Engineer (GPU/CUDA)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Equity
Medical benefits (full coverage)
PTO & Hybrid work policy
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
High-Performance ML Inference Engineer
High-Performance ML Inference Engineer

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
Low-Latency Inference Infrastructure Engineer
Low-Latency Inference Infrastructure Engineer

Elorian • Palo Alto (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior ML Infra Engineer - Low-Latency Scale
Senior ML Infra Engineer - Low-Latency Scale

Patreon • New York (NY), San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Healthcare
401k with matching
Paid time off