Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP

Bala Cynwyd (PA)

On-site

USD 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Susquehanna International Group, LLP is seeking a Machine Learning Engineer in Bala Cynwyd, PA. This role focuses on low-latency inference optimization for high-performance model serving systems.

You will collaborate with researchers to optimize performance, evaluate frameworks, and debug GPU memory issues while managing inference workloads effectively. A strong background in modern ML frameworks, programming experience, and understanding of production environments is essential.

Qualifications

  • Experience deploying, optimizing machine learning inference workloads in production.
  • Programming experience in Python, Java, C#, and systems languages like C or C++.
  • Strong understanding of modern ML frameworks like PyTorch.

Responsibilities

  • Design and optimize low-latency inference systems for production ML workloads.
  • Profile model inference pipelines for performance improvements.
  • Debug performance issues related to GPU and CPU coordination.

Skills

Machine Learning inference optimization
Python
C++
PyTorch
Kubernetes

Education

Background in mathematics, physics, or computer science

Tools

CUDA
Triton
GPU clusters

Job description

Overview

We are looking for a Machine Learning Engineer focused on low-latency inference optimization to help build, tune, and productionize high-performance model serving systems. This role sits at the intersection of machine learning, systems engineering, and GPU performance. You will work on inference workloads where latency, throughput, reliability, and hardware efficiency all matter, and where a deep understanding of modern inference runtimes can meaningfully improve production outcomes.

You will work closely with quantitative researchers and engineers to understand model structure, identify inference bottlenecks, and turn research ideas into efficient production systems. The work may involve other types of models, but focuses on transformer-style architectures, and structured inference workloads. You will evaluate and tune frameworks and related serving or compilation systems, while also reasoning about GPU execution, memory layout, batching strategies, precision tradeoffs, and end-to-end latency.

What you'll do
  • Design, build, and optimize low-latency inference systems for production machine learning workloads.
  • Profile model inference pipelines across model execution, runtime configuration, batching, memory movement, serialization, networking, and I/O.
  • Evaluate, integrate, and tune inference runtime systems.
  • Improve latency, throughput, GPU utilization, for production inference workloads.
  • Build and support benchmarking and profiling tools to compare model variants, hardware targets, runtime configurations, and deployment strategies.
  • Debug performance issues involving GPU memory, compute saturation, kernel behavior, CPU/GPU coordination, data movement, and serving-layer overhead.
  • Help shape model and system design choices so that research models are efficient to deploy under real latency constraints.
  • Where necessary, collaborate with lower-level systems or GPU specialists on custom operators, kernel-level optimization, or hardware-specific performance work.
What we’re looking for
  • Experience deploying, optimizing, or operating machine learning inference workloads in production or production-like environments.
  • Programming experience in Python, Java, C# etc. and at least one systems language such as C, C++, Rust, or Go
  • Solid understanding of modern ML frameworks such as PyTorch, including model execution, export, tracing, compilation, and performance profiling.
  • Ability to reason about latency, throughput, batching, memory use, GPU utilization, and reliability under real workloads.
  • Strong practical judgment around tradeoffs between model quality, latency, throughput, implementation complexity, and maintainability.
Preferred qualifications
  • Experience optimizing inference for latency-sensitive or high-throughput applications.
  • Experience with model optimization techniques such as quantization, pruning, distillation, operator fusion, graph lowering, custom operators, or model compilation.
  • Exposure to CUDA, Triton language, ROCm, PTX, CuTe, CUTLASS, FlashInfer, or similar low-level GPU programming tools.
  • Experience running inference workloads on Kubernetes or GPU clusters, including scheduling, autoscaling, observability, and resource management.
  • Background in mathematics, physics, computer science, engineering, statistics, quantitative finance, or another technical field.
  • Demonstrated ability to improve real-world inference performance beyond a baseline framework implementation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Inference Optimization ML Engineer
Inference Optimization ML Engineer

Rhoda AI • Mountain View (WY)

On-site
USD 180,000 - 260,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Staff Machine Learning Software Engineer
Senior Staff Machine Learning Software Engineer

San Diego Stealth Startup • San Diego (CA)

On-site
USD 202,000 - 215,000
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000