Low-Latency ML Inference Engineer

Career Techniques

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Career Techniques in New York, NY seeks a hands-on ML inference engineer to optimize production-grade models across CPUs, GPUs, and FPGAs. You will lead platform evaluations, tune kernels, and push memory and interconnect performance to meet latency requirements.

Collaborate with ML researchers, HPC and datacenter teams to deploy compact, low-latency inference workloads, using tools like Triton, TensorRT and Nsight, with a focus on reliability and scalable throughput.

Qualifications

  • 2+ years optimizing DL inference in latency-sensitive environments.
  • Deep expertise in PyTorch/JAX with Python/C++.
  • Experience in custom GPU kernel development and libraries for performance.
  • Strong knowledge of GPU architecture and memory hierarchies.
  • Experience benchmarking across heterogeneous compute architectures.

Responsibilities

  • Lead evaluation of inference platforms across CPUs, GPUs, and FPGAs.
  • Analyze and optimize execution across deep memory hierarchies.
  • Collaborate with infrastructure teams on latency-critical strategy.
  • Develop highly optimized GPU kernels and integrate libraries.
  • Implement model compression for low-latency inference.
  • Work with ML researchers, HPC/FPGA/datacenter teams on deployments.

Skills

PyTorch
JAX
Python
C++
Mixed-precision
GPU kernel development
TensorRT
Profiling tools
CuBLAS/CUTLASS
FPGAs/ASICs

Tools

Triton
TensorRT
ONNX
IREE
HLS4ML
cuBLAS
CUTLASS
Nsight Systems
Nsight Compute

Job description

Career Techniques in New York, NY seeks a hands-on ML inference engineer to optimize production-grade models across CPUs, GPUs, and FPGAs. You will lead platform evaluations, tune kernels, and push memory and interconnect performance to meet latency requirements.

Collaborate with ML researchers, HPC and datacenter teams to deploy compact, low-latency inference workloads, using tools like Triton, TensorRT and Nsight, with a focus on reliability and scalable throughput.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid Low-Latency ML Inference Engineer
Hybrid Low-Latency ML Inference Engineer

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Free breakfast, lunch, and snacks
Wellness expense reimbursement
+3
Low-Latency ML Inference Engineer (GPU/FPGA)
Low-Latency ML Inference Engineer (GPU/FPGA)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous paid time off policies
Hybrid working opportunities
In-office wellness experiences
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
ML Inference Engineer - Low-Latency GPU Systems
ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Senior ML Infra Engineer: Low-Latency Inference
Senior ML Infra Engineer: Low-Latency Inference

Doist • New York (NY)

Hybrid
USD 170,000 - 250,000
Healthcare
401k plan with matching
Hybrid work model
+1
Senior ML Systems Engineer: Low-Latency Inference
Senior ML Systems Engineer: Low-Latency Inference

Strativ Group • Palo Alto (CA)

On-site
USD 500,000 - 600,000
ML Systems Engineer: Scalable Training & Realtime Inference
ML Systems Engineer: Scalable Training & Realtime Inference

Jobzhr • New York (NY)

On-site
USD 180,000 - 280,000