CUDA Kernel Architect for Low-Latency GPU Inference

Trading Interview

Northern (KY)

Hybrid

USD 140,000 - 210,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Susquehanna is seeking a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role targets workloads where standard runtimes fall short, demanding custom kernels, memory layouts, and execution strategies to gain meaningful performance.

You will partner with quantitative researchers to identify bottlenecks, translate models into efficient GPU implementations, and push end-to-end latency improvements in production systems.

Qualifications

  • Expert in writing and optimizing CUDA kernels.
  • Strong C/C++ programming experience.
  • Deep understanding of GPU architecture and memory hierarchy.
  • Ability to reason about numerical stability and hardware efficiency.
  • Experience with low-level systems and performance analysis.

Responsibilities

  • Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads.
  • Develop fine-grained GPU implementations tailored to model structures.
  • Analyze models to identify bottlenecks and opportunities for parallelization.
  • Collaborate with quantitative researchers to translate models into high-performance pipelines.
  • Profile, benchmark, and improve end-to-end inference latency and throughput.

Skills

CUDA kernels
C/C++
GPU architecture
numerical stability
low-level systems

Job description

Susquehanna is seeking a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role targets workloads where standard runtimes fall short, demanding custom kernels, memory layouts, and execution strategies to gain meaningful performance.

You will partner with quantitative researchers to identify bottlenecks, translate models into efficient GPU implementations, and push end-to-end latency improvements in production systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Performance Architect: Low-Latency CUDA Mastery
GPU Performance Architect: Low-Latency CUDA Mastery

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Low-Latency GPU Kernel Engineer
Low-Latency GPU Kernel Engineer

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
Latency-Critical CUDA Kernel Engineer
Latency-Critical CUDA Kernel Engineer

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Trading Interview • Northern (KY)

Hybrid
USD 140,000 - 210,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Senior Inference Engineer: GPU Kernel Optimization Equity
Senior Inference Engineer: GPU Kernel Optimization Equity

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 184,000 - 288,000
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
ML Inference Engineer - Low-Latency GPU Systems
ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000