GPU Developer

SIG

Hong Kong

On-site

HKD 900,000 - 1,300,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Susquehanna is seeking a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role targets workloads where standard libraries don’t fully exploit model structures, enabling custom kernels and memory layouts to deliver meaningful gains.

You will partner with quantitative researchers to identify bottlenecks, translate math into production-grade GPU code, and push latency and throughput to meet demanding trading workloads.

Qualifications

  • Proven ability to write and optimize CUDA kernels.
  • Strong C/C++ programming experience.
  • Deep understanding of GPU architecture and memory hierarchy.
  • Ability to reason about numerical stability and performance tradeoffs.
  • Experience with low-level systems.

Responsibilities

  • Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads.
  • Develop fine-grained GPU implementations tailored to specific model structures.
  • Analyze models to identify bottlenecks and opportunities for parallelization.
  • Collaborate with researchers to translate models into high-performance pipelines.
  • Optimize end-to-end inference performance through tuning and memory layouts.
  • Profile and benchmark GPU performance.
  • Improve latency and throughput in production systems.
  • Contribute to GPU architecture decisions and performance best practices.

Skills

CUDA kernels
C/C++
GPU architecture
Numerical stability
Low-level systems

Education

PhD in Mathematics, Physics, Computer Science, Engineering, or related quantitative field

Tools

ONNX Runtime
TensorRT
Triton
TVM
PTX-level behavior

Job description

Overview

We are looking for a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role is focused on workloads where off-the-shelf runtimes and vendor libraries do not fully exploit the structure of the model, and where custom kernels, memory layouts, and execution strategies can deliver meaningful gains.

You will work closely with quantitative researchers and engineers to understand model structure, identify computational bottlenecks, and turn mathematical ideas into production-grade GPU implementations. You will use your understanding of GPU hardware to help shape models that are both mathematically effective and efficient to run. The problems span compact neural networks, tree-based models, and other structured inference workloads where latency, throughput, and efficiency all matter.

This role is a strong fit for someone who enjoys low-level optimization, performance analysis, and translating abstract models into hardware-efficient code.

Key Responsibilities
  • Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads
  • Develop fine-grained GPU implementations tailored to specific model structures
  • Analyze quantitative research models and computational bottlenecks to identify opportunities for parallelization and hardware-efficient execution
  • Collaborate directly with quantitative researchers to translate mathematical models into high-performance compute pipelines
  • Optimize end-to-end inference performance through kernel tuning, memory-layout design, execution strategy, I/O optimization, and precision tradeoffs
  • Profile and benchmark GPU performance
  • Improve latency and throughput in production inference systems
  • Contribute to GPU architecture decisions and performance best practices
What you can expect from us

Real Impact:By integrating sophisticated coding techniques with innovative engineering ideas, we design and optimize systems that can process massive amounts of data while still ensuring high performance and stability. You'll see how your contributions towards developing and supporting leading-edge hardware and software technologiesmake a firm-wide impact that makes us all smarter, faster, and better.

Collaboration: You will partner closely with our Infrastructure and Strategy Developer Teams to deliver and manage scalable and highly performant systems to support trading in the region and globally.

Growth: For many of our roles, we don’t expect you to have prior industry experience in proprietary trading or financial services to succeed at Susquehanna. We're looking for people who are naturally curious, relentless problem solvers, and have the desire to continuously innovate, learn, and grow.

Benefits: Susquehanna offers a wide array of competitive employee perks & benefits

What we're looking for
  • Strong proficiency in writing and optimizing CUDA kernels
  • Solid programming experience in C/C++ (preferred)
  • Deep understanding of GPU architecture, including memory hierarchy, SIMT execution, occupancy, and latency/throughput tradeoffs
  • Ability to reason about numerical stability, precision, performance tradeoffs, and how model design choices affect hardware efficiency
  • Strong problem-solving skills and comfort working with low-level systems
Preferred Qualifications
  • PhD in Mathematics, Physics, Computer Science, Engineering, or related quantitative field
  • Strong background in linear algebra, probability, numerical methods, or scientific computing
  • Experience working with quantitative research teams or financial models
  • Demonstrated ability to improve real-world inference performance beyond baseline framework or library implementations
  • Familiarity with PTX-level behavior, tensor core utilization, or architecture-specific tuning
  • Exposure to ONNX Runtime, TensorRT, Triton, TVM, or similar systems
  • Exposure to:
  • Neural Networks
  • Tree-based models (e.g. KightGBM)
  • State space models (e.g. Mamba architectures)
  • Experience with kernel fusion, custom operators, model compilation, or graph-level optimization
About Susquehanna

Susquehanna is a global quantitative trading firm founded by a group of friends who share a passion for game theory and probabilistic thinking. We have incorporated this approach into our culture, where you will find relentless problem solvers within each of our core disciplines: Trading, Technology, and Quantitative Research. From offices around the world, our employees collaborate to make optimal decisions and are driven by the desire to achieve winning results together.

What we do

We are experts in trading essentially all listed financial products and asset classes, with a focus on derivatives trading. Through market making and market taking, we handle millions of trading transactions around the world every day, providing liquidity and ensuring competitive prices for buyers and sellers. While our presence in the market is broad, our trading desks are highly specialised, allowing for a deep understanding of unique drivers of each asset class.

Equal Opportunity Statement

We encourage applications from candidates from all backgrounds, and we welcome requests for reasonableadjustmentsduring the recruitment process to ensure that you can best demonstrate your abilities.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Developer
GPU Developer

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 900,000 - 1,300,000
CUDA Kernel Engineer for Low-Latency Inference
CUDA Kernel Engineer for Low-Latency Inference

SIG • Hong Kong

On-site
HKD 900,000 - 1,300,000
Operations Tech Analyst
Operations Tech Analyst

Susquehanna International Group • Hong Kong

On-site
HKD 450,000 - 700,000
Health insurance
Lunch allowance
Education allowances
+1
Operations Tech Analyst
Operations Tech Analyst

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 480,000 - 720,000
Health insurance
Education allowances
Daily lunch allowance
Machine Learning Engineering Internship: Summer 2027
Machine Learning Engineering Internship: Summer 2027

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 167,000 - 279,000
CUDA Kernel Architect for Low-Latency Inference
CUDA Kernel Architect for Low-Latency Inference

Susquehanna International Group, LLP • Hong Kong

On-site
HKD 900,000 - 1,300,000
Security Engineer
Security Engineer

Susquehanna International Group • Hong Kong

On-site
HKD 420,000 - 720,000
Quantitative Strategy Developer Internship: Summer 2027
Quantitative Strategy Developer Internship: Summer 2027

Susquehanna International Group • Hong Kong

On-site
HKD 349,843 - 524,764
Travel and housing covered
Fully stocked kitchens
On-site Wellness Center
+2
Application Support Engineer
Application Support Engineer

Susquehanna International Group • Hong Kong

On-site
HKD 420,000 - 540,000
Quantitative Systematic Trading Internship - Master's: Summer 2027
Quantitative Systematic Trading Internship - Master's: Summer 2027

Susquehanna International Group • Hong Kong

On-site
HKD 200,000 - 300,000