CUDA Kernel Architect for Low-Latency Inference

Susquehanna International Group, LLP

Hong Kong

On-site

HKD 900,000 - 1,300,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Susquehanna International Group, LLP is seeking a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. You’ll work with quantitative researchers to identify bottlenecks and translate mathematical models into production-grade GPU code, focusing on models like compact neural networks and tree-based methods where latency and throughput matter.

You should enjoy low-level optimization, performance analysis, and hardware-aware design.

Qualifications

  • Strong proficiency in writing and optimizing CUDA kernels.
  • Solid programming experience in C/C++ (preferred).
  • Deep understanding of GPU architecture, including memory hierarchy, SIMT execution, occupancy, and latency/throughput tradeoffs.
  • Ability to reason about numerical stability, precision, performance tradeoffs, and how model design choices affect hardware efficiency.
  • Strong problem-solving skills and comfort working with low-level systems.

Responsibilities

  • Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads
  • Develop fine-grained GPU implementations tailored to specific model structures
  • Analyze quantitative research models and computational bottlenecks to identify opportunities for parallelization and hardware-efficient execution
  • Collaborate directly with quantitative researchers to translate mathematical models into high-performance compute pipelines
  • Optimize end-to-end inference performance through kernel tuning, memory-layout design, execution strategy, I/O optimization, and precision tradeoffs
  • Profile and benchmark GPU performance
  • Improve latency and throughput in production inference systems
  • Contribute to GPU architecture decisions and performance best practices

Skills

CUDA kernels
C/C++
GPU architecture
Numerical stability
Low-level systems

Education

PhD in Mathematics, Physics, Computer Science, Engineering, or related quantitative field

Tools

ONNX Runtime
TensorRT
Triton
TVM

Job description

Susquehanna International Group, LLP is seeking a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. You’ll work with quantitative researchers to identify bottlenecks and translate mathematical models into production-grade GPU code, focusing on models like compact neural networks and tree-based methods where latency and throughput matter.

You should enjoy low-level optimization, performance analysis, and hardware-aware design.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

CUDA Kernel Engineer for Low-Latency Inference
CUDA Kernel Engineer for Low-Latency Inference

SIG • Hong Kong

On-site
HKD 900,000 - 1,300,000
GPU Developer
GPU Developer

SIG • Hong Kong

On-site
HKD 900,000 - 1,300,000
GPU Developer
GPU Developer

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 900,000 - 1,300,000
Inference & System Optimization Engineer (Experienced )
Inference & System Optimization Engineer (Experienced )

奇瑞全球創新(香港)有限公司 • Hong Kong Island

On-site
HKD 900,000 - 1,300,000
Edge Inference & Systems Optimization Engineer
Edge Inference & Systems Optimization Engineer

奇瑞全球創新(香港)有限公司 • Hong Kong Island

On-site
HKD 900,000 - 1,300,000
Principal GPU Memory Systems Engineer — Scale-Up/Out
Principal GPU Memory Systems Engineer — Scale-Up/Out

METAVERSE COMPUTING LIMITED • Hong Kong

On-site
HKD 1,200,000 - 1,600,000
Principal engineer – GPU Memory Systems & Scale-Up/Out Systems
Principal engineer – GPU Memory Systems & Scale-Up/Out Systems

METAVERSE COMPUTING LIMITED • Hong Kong

On-site
HKD 1,200,000 - 1,600,000
Low-Latency C++ Quant Developer for HFT
Low-Latency C++ Quant Developer for HFT

Tower Research Capital • Hong Kong

Hybrid
HKD 500,000 - 1,000,000
Hybrid work
Free meals
Wellness reimbursements
+3
Distributed ML Engineer: Real-Time Inference & GPU Training
Distributed ML Engineer: Real-Time Inference & GPU Training

IMC Trading • Hong Kong

On-site
HKD 700,000 - 1,100,000
Low-Latency C++ Engineer for Global Quant Trading
Low-Latency C++ Engineer for Global Quant Trading

Pinpoint Asia • Hong Kong

On-site
HKD 900,000 - 1,300,000