ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP

Bala Cynwyd (PA)

On-site

USD 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Susquehanna International Group, LLP is seeking a Machine Learning Engineer in Bala Cynwyd, PA. This role focuses on low-latency inference optimization for high-performance model serving systems.

You will collaborate with researchers to optimize performance, evaluate frameworks, and debug GPU memory issues while managing inference workloads effectively. A strong background in modern ML frameworks, programming experience, and understanding of production environments is essential.

Qualifications

  • Experience deploying, optimizing machine learning inference workloads in production.
  • Programming experience in Python, Java, C#, and systems languages like C or C++.
  • Strong understanding of modern ML frameworks like PyTorch.

Responsibilities

  • Design and optimize low-latency inference systems for production ML workloads.
  • Profile model inference pipelines for performance improvements.
  • Debug performance issues related to GPU and CPU coordination.

Skills

Machine Learning inference optimization
Python
C++
PyTorch
Kubernetes

Education

Background in mathematics, physics, or computer science

Tools

CUDA
Triton
GPU clusters

Job description

Susquehanna International Group, LLP is seeking a Machine Learning Engineer in Bala Cynwyd, PA. This role focuses on low-latency inference optimization for high-performance model serving systems.

You will collaborate with researchers to optimize performance, evaluate frameworks, and debug GPU memory issues while managing inference workloads effectively. A strong background in modern ML frameworks, programming experience, and understanding of production environments is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Low-Latency GPU Kernel Engineer
Low-Latency GPU Kernel Engineer

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Latency-Critical CUDA Kernel Engineer
Latency-Critical CUDA Kernel Engineer

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Architect: Low-Latency CUDA Mastery
GPU Performance Architect: Low-Latency CUDA Mastery

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
CUDA Kernel Architect for Low-Latency GPU Inference
CUDA Kernel Architect for Low-Latency GPU Inference

Trading Interview • Northern (KY)

Hybrid
USD 140,000 - 210,000
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
ML Performance Engineer: Low-Level Systems & GPUs
ML Performance Engineer: Low-Level Systems & GPUs

Trading Interview • New York (NY)

On-site
USD 170,000 - 210,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000