ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP

Bala Cynwyd (PA)

On-site

USD 110,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Susquehanna International Group, LLP is seeking a Machine Learning Engineer in Bala Cynwyd, PA. This role focuses on low-latency inference optimization for high-performance model serving systems.

You will collaborate with researchers to optimize performance, evaluate frameworks, and debug GPU memory issues while managing inference workloads effectively. A strong background in modern ML frameworks, programming experience, and understanding of production environments is essential.

Qualifications

  • Experience deploying, optimizing machine learning inference workloads in production.
  • Programming experience in Python, Java, C#, and systems languages like C or C++.
  • Strong understanding of modern ML frameworks like PyTorch.

Responsibilities

  • Design and optimize low-latency inference systems for production ML workloads.
  • Profile model inference pipelines for performance improvements.
  • Debug performance issues related to GPU and CPU coordination.

Skills

Machine Learning inference optimization
Python
C++
PyTorch
Kubernetes

Education

Background in mathematics, physics, or computer science

Tools

CUDA
Triton
GPU clusters

Job description

Susquehanna International Group, LLP is seeking a Machine Learning Engineer in Bala Cynwyd, PA. This role focuses on low-latency inference optimization for high-performance model serving systems.

You will collaborate with researchers to optimize performance, evaluate frameworks, and debug GPU memory issues while managing inference workloads effectively. A strong background in modern ML frameworks, programming experience, and understanding of production environments is essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
GPU Performance Architect: Low-Latency CUDA Mastery
GPU Performance Architect: Low-Latency CUDA Mastery

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Latency-Critical CUDA Kernel Engineer
Latency-Critical CUDA Kernel Engineer

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
CUDA Kernel Architect for Low-Latency GPU Inference
CUDA Kernel Architect for Low-Latency GPU Inference

Trading Interview • Northern (KY)

Hybrid
USD 140,000 - 210,000
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
ML Inference Systems Engineer – Remote, Scalable, Low-Latency
ML Inference Systems Engineer – Remote, Scalable, Low-Latency

Atlassian Corp. • Seattle (WA), Northern (KY)

Hybrid
USD 178,000 - 233,000
Health and wellbeing resources
Volunteer days
LLM Inference Optimization Engineer - Frontier Performance
LLM Inference Optimization Engineer - Frontier Performance

GMI Cloud, Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000
Microsecond ML Inference Architect
Microsecond ML Inference Architect

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Generous PTO
Hybrid work options
Free breakfast, lunch, snacks
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
ML Performance Engineer: Low-Level Systems & GPUs
ML Performance Engineer: Low-Level Systems & GPUs

Trading Interview • New York (NY)

On-site
USD 170,000 - 210,000