Machine Learning Performance Engineer (Inference)

Career Techniques

New York (NY)

On-site

USD 200,000 - 300,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Career Techniques in New York, NY seeks a hands-on ML inference engineer to optimize production-grade models across CPUs, GPUs, and FPGAs. You will lead platform evaluations, tune kernels, and push memory and interconnect performance to meet latency requirements.

Collaborate with ML researchers, HPC and datacenter teams to deploy compact, low-latency inference workloads, using tools like Triton, TensorRT and Nsight, with a focus on reliability and scalable throughput.

Qualifications

  • 2+ years optimizing DL inference in latency-sensitive environments.
  • Deep expertise in PyTorch/JAX with Python/C++.
  • Experience in custom GPU kernel development and libraries for performance.
  • Strong knowledge of GPU architecture and memory hierarchies.
  • Experience benchmarking across heterogeneous compute architectures.

Responsibilities

  • Lead evaluation of inference platforms across CPUs, GPUs, and FPGAs.
  • Analyze and optimize execution across deep memory hierarchies.
  • Collaborate with infrastructure teams on latency-critical strategy.
  • Develop highly optimized GPU kernels and integrate libraries.
  • Implement model compression for low-latency inference.
  • Work with ML researchers, HPC/FPGA/datacenter teams on deployments.

Skills

PyTorch
JAX
Python
C++
Mixed-precision
GPU kernel development
TensorRT
Profiling tools
CuBLAS/CUTLASS
FPGAs/ASICs

Tools

Triton
TensorRT
ONNX
IREE
HLS4ML
cuBLAS
CUTLASS
Nsight Systems
Nsight Compute

Job description

Responsibilities
  • Benchmarking & Strategy:
    • Lead the technical evaluation of diverse inference platforms - ranging across CPUs, GPUs, and FPGAs - to guide infrastructure deployment decisions.
  • System Architecture Optimization:
    • Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing. You will assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle.
  • Infrastructure & Deployment Feasibility:
    • Collaborate with Infrastructure teams to understand thermal, power, and operational constraints of hardware platforms to design inference strategies for our latency-critical trading strategies that fit within those envelopes.
  • GPU Kernel Development:
    • Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon.
  • Model Optimization & Deployment:
    • Implement advanced model reduction techniques (quantization, pruning, distillation) to ensure compact memory footprints and numerical stability. Prioritize optimization for low-latency, event-level inference workloads to meet real-time trading requirements.
  • Cross-Functional Collaboration:
    • Collaborate closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring to fruition target deployments.
Qualifications
  • 2+ years of experience optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain.
  • ML Frameworks: Deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation.
  • Kernel Development & Optimization Tooling: Proven experience in custom GPU kernel development. Deep familiarity with advanced optimization libraries and compilers (e.g., Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) as well as profiling tools (e.g., Nsight Systems, Nsight Compute).
  • GPU Architecture Mastery: Deep expertise in GPU microarchitecture, encompassing SM execution, warp scheduling, and full memory hierarchy optimization (registers to HBM).
  • Cross-Architecture Benchmarking: Proven record of rigorous, data-driven approach to evaluating inference performance across heterogeneous compute architectures.
  • Bonus: Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs.
  • Prior experience in financial trading is not required.

Comp: $200-300K + Bonus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Machine Learning Engineer (Training & Inference Systems)
Machine Learning Engineer (Training & Inference Systems)

Fintal Partners • New York (NY)

On-site
USD 185,000 - 230,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Selby Jennings • Chicago (IL)

On-site
USD 140,000 - 210,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Optiver • New York (NY)

On-site
USD 200,000
Global profit-sharing pool
401(k) match up to 50%
Comprehensive health coverage
+2
Senior ML Performance Engineer
Senior ML Performance Engineer

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
Machine Learning Performance Engineer - Quant Research & Trading
Machine Learning Performance Engineer - Quant Research & Trading

Acquire Me • United States

On-site
USD 200,000 - 350,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits