Latency-Critical CUDA Kernel Engineer

Susquehanna International Group

Bala Cynwyd (PA)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Susquehanna International Group is seeking a GPU Performance Engineer in Bala Cynwyd, Pennsylvania. In this role, you will design and optimize CUDA kernels for low-latency inference workloads, working closely with quantitative researchers. This position requires strong expertise in GPU architecture and programming in C/C++. A PhD in a quantitative field is preferred. The role focuses on enhancing performance for various structured inference workloads and demands a solid understanding of GPU hardware and optimization techniques.

Qualifications

  • Strong proficiency in writing and optimizing CUDA kernels.
  • Deep understanding of GPU architecture.
  • Experience with quantitative research teams or financial models.

Responsibilities

  • Design, implement, and optimize custom CUDA kernels.
  • Analyze research models to identify opportunities for optimization.
  • Collaborate with quantitative researchers to translate models into compute pipelines.

Skills

Proficiency in CUDA
C/C++ programming
Problem-solving skills
GPU architecture knowledge

Education

PhD in mathematics, physics, computer science, engineering, or related quantitative field

Tools

ONNX Runtime
TensorRT

Job description

Susquehanna International Group is seeking a GPU Performance Engineer in Bala Cynwyd, Pennsylvania. In this role, you will design and optimize CUDA kernels for low-latency inference workloads, working closely with quantitative researchers. This position requires strong expertise in GPU architecture and programming in C/C++. A PhD in a quantitative field is preferred. The role focuses on enhancing performance for various structured inference workloads and demands a solid understanding of GPU hardware and optimization techniques.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Kernel Architect for Low-Latency GPU Inference
CUDA Kernel Architect for Low-Latency GPU Inference

Trading Interview • Northern (KY)

Hybrid
USD 140,000 - 210,000
GPU Performance Architect: Low-Latency CUDA Mastery
GPU Performance Architect: Low-Latency CUDA Mastery

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
ML Inference Engineer - Low-Latency GPU Systems
ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Trading Interview • Northern (KY)

Hybrid
USD 140,000 - 210,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Kernel Engineer: Build Fast AI Inference at Scale
GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
Senior GPU Kernel Performance Engineer - Hybrid, Bellevue
Senior GPU Kernel Performance Engineer - Hybrid, Bellevue

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan
Annual bonus
+1
GPU Kernel Engineer for AI Inference & Performance
GPU Kernel Engineer for AI Inference & Performance

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
GPU Compute Engineer — HPC & DL Inference Optimizer
GPU Compute Engineer — HPC & DL Inference Optimizer

Luxoft Poland • Town of Poland (NY)

On-site
USD 110,000 - 170,000
Stable employment
LuxMed health care
Life and travel insurance
+5