CUDA Kernel Architect for Low-Latency GPU Inference

Trading Interview

Northern (KY)

Hybrid

USD 140,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Susquehanna is seeking a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role targets workloads where standard runtimes fall short, demanding custom kernels, memory layouts, and execution strategies to gain meaningful performance.

You will partner with quantitative researchers to identify bottlenecks, translate models into efficient GPU implementations, and push end-to-end latency improvements in production systems.

Qualifications

  • Expert in writing and optimizing CUDA kernels.
  • Strong C/C++ programming experience.
  • Deep understanding of GPU architecture and memory hierarchy.
  • Ability to reason about numerical stability and hardware efficiency.
  • Experience with low-level systems and performance analysis.

Responsibilities

  • Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads.
  • Develop fine-grained GPU implementations tailored to model structures.
  • Analyze models to identify bottlenecks and opportunities for parallelization.
  • Collaborate with quantitative researchers to translate models into high-performance pipelines.
  • Profile, benchmark, and improve end-to-end inference latency and throughput.

Skills

CUDA kernels
C/C++
GPU architecture
numerical stability
low-level systems

Job description

Susquehanna is seeking a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role targets workloads where standard runtimes fall short, demanding custom kernels, memory layouts, and execution strategies to gain meaningful performance.

You will partner with quantitative researchers to identify bottlenecks, translate models into efficient GPU implementations, and push end-to-end latency improvements in production systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Performance Architect: Low-Latency CUDA Mastery
GPU Performance Architect: Low-Latency CUDA Mastery

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Latency-Critical CUDA Kernel Engineer
Latency-Critical CUDA Kernel Engineer

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Trading Interview • Northern (KY)

On-site
USD 140,000 - 210,000
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
ML Inference Engineer - Low-Latency GPU Systems
ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Obsidian • San Francisco (CA)

On-site
USD 80,000 - 120,000
Inference Runtime Performance Engineer — GPU Kernels
Inference Runtime Performance Engineer — GPU Kernels

iFrame Corporation • San Francisco (CA)

Remote
USD 220,000 - 360,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000