CUDA Kernel Engineer for Low-Latency Inference

SIG

Hong Kong

On-site

HKD 900,000 - 1,300,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Susquehanna is seeking a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role targets workloads where standard libraries don’t fully exploit model structures, enabling custom kernels and memory layouts to deliver meaningful gains.

You will partner with quantitative researchers to identify bottlenecks, translate math into production-grade GPU code, and push latency and throughput to meet demanding trading workloads.

Qualifications

  • Proven ability to write and optimize CUDA kernels.
  • Strong C/C++ programming experience.
  • Deep understanding of GPU architecture and memory hierarchy.
  • Ability to reason about numerical stability and performance tradeoffs.
  • Experience with low-level systems.

Responsibilities

  • Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads.
  • Develop fine-grained GPU implementations tailored to specific model structures.
  • Analyze models to identify bottlenecks and opportunities for parallelization.
  • Collaborate with researchers to translate models into high-performance pipelines.
  • Optimize end-to-end inference performance through tuning and memory layouts.
  • Profile and benchmark GPU performance.
  • Improve latency and throughput in production systems.
  • Contribute to GPU architecture decisions and performance best practices.

Skills

CUDA kernels
C/C++
GPU architecture
Numerical stability
Low-level systems

Education

PhD in Mathematics, Physics, Computer Science, Engineering, or related quantitative field

Tools

ONNX Runtime
TensorRT
Triton
TVM
PTX-level behavior

Job description

Susquehanna is seeking a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role targets workloads where standard libraries don’t fully exploit model structures, enabling custom kernels and memory layouts to deliver meaningful gains.

You will partner with quantitative researchers to identify bottlenecks, translate math into production-grade GPU code, and push latency and throughput to meet demanding trading workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

CUDA Kernel Architect for Low-Latency Inference
CUDA Kernel Architect for Low-Latency Inference

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 900,000 - 1,300,000
GPU Developer
GPU Developer

SIG • Hong Kong

On-site
HKD 900,000 - 1,300,000
GPU Developer
GPU Developer

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 900,000 - 1,300,000
Inference & System Optimization Engineer (Experienced )
Inference & System Optimization Engineer (Experienced )

奇瑞全球創新(香港)有限公司 • Hong Kong Island

On-site
HKD 900,000 - 1,300,000
Low-Latency C++ Quant Developer for HFT
Low-Latency C++ Quant Developer for HFT

Tower Research Capital • Hong Kong

Hybrid
HKD 500,000 - 1,000,000
Hybrid work
Free meals
Wellness reimbursements
+3
Low-Latency C++ Engineer for Global Quant Trading
Low-Latency C++ Engineer for Global Quant Trading

Pinpoint Asia • Hong Kong

On-site
HKD 900,000 - 1,300,000
Ultra-Low Latency C++ Engineer for Real-Time Trading
Ultra-Low Latency C++ Engineer for Real-Time Trading

Selby Jennings • Hong Kong

On-site
HKD 1,200,000 - 1,600,000
Distributed ML Engineer: Real-Time Inference & GPU Training
Distributed ML Engineer: Real-Time Inference & GPU Training

IMC Trading • Hong Kong

On-site
HKD 700,000 - 1,100,000
Edge Inference & Systems Optimization Engineer
Edge Inference & Systems Optimization Engineer

奇瑞全球創新(香港)有限公司 • Hong Kong Island

On-site
HKD 900,000 - 1,300,000
Machine Learning Engineer
Machine Learning Engineer

ittihad medical centre • Hong Kong

On-site
HKD 500,000 - 700,000