GPU Performance Engineer | Experienced Hire

SIG Susquehanna

Bala Cynwyd (PA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SIG Susquehanna in Bala Cynwyd, Pennsylvania, is seeking a GPU Performance Engineer to develop optimized CUDA kernels for low-latency inference. The ideal candidate has strong CUDA programming skills and expertise in GPU architecture. Responsibilities include optimizing kernel performance, analyzing models for computational bottlenecks, and enhancing production inference systems. A PhD in a relevant field is preferred. Join an engaging environment committed to performance improvements and innovation in computational efficiency.

Qualifications

  • Strong proficiency in writing and optimizing CUDA kernels.
  • Solid programming experience in C/C++.
  • Deep understanding of GPU architecture and performance tradeoffs.
  • Ability to reason about numerical stability and precision.
  • Strong problem-solving skills.

Responsibilities

  • Design, implement, and optimize custom CUDA kernels.
  • Develop GPU implementations tailored to specific model structures.
  • Analyze computational bottlenecks and identify optimization opportunities.
  • Collaborate with researchers to translate models into high-performance pipelines.
  • Improve end-to-end inference performance.

Skills

CUDA optimization
C/C++ programming
GPU architecture understanding
Problem-solving in low-level systems

Education

PhD in mathematics, physics, computer science, engineering, or related field

Tools

ONNX Runtime
TensorRT
Triton
TVM

Job description

Overview

We are looking for a GPU Performance Engineer to build highly optimized CUDA kernels for low-latency inference. This role focuses on workloads where off-the-shelf runtimes and vendor libraries do not fully exploit the structure of the model, and where custom kernels, memory layouts, and execution strategies can deliver meaningful gains.

You will work closely with quantitative researchers and engineers to understand model structure, identify computational bottlenecks, and convert mathematical ideas into production-grade GPU implementations. Using your understanding of GPU hardware, you will help shape models that are both mathematically effective and efficient to run. The problems span compact neural networks, tree-based models, and other structured inference workloads where latency, throughput, and efficiency all matter.

This role is a strong fit for someone who enjoys low-level optimization, performance analysis, and translating abstract models into hardware-efficient code.

What you’ll do
  • Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads
  • Develop fine-grained GPU implementations tailored to specific model structures
  • Analyze quantitative research models and computational bottlenecks to identify opportunities for parallelization and hardware-efficient execution
  • Collaborate directly with quantitative researchers to translate mathematical models into high-performance computing pipelines
  • Optimize end-to-end inference performance through kernel tuning, memory‑layout design, execution strategy, I/O optimization, and precision tradeoffs
  • Profile and benchmark GPU performance
  • Improve latency and throughput in production inference systems
  • Contribute to GPU architecture decisions and performance best practices
What we’re looking for
  • Strong proficiency in writing and optimizing CUDA kernels
  • Solid programming experience in C/C++ (preferred)
  • Deep understanding of GPU architecture, including memory hierarchy, SIMT execution, occupancy, and latency/throughput tradeoffs
  • Ability to reason about numerical stability, precision, performance tradeoffs, and how model design choices affect hardware efficiency
  • Strong problem‑solving skills and comfort working with low-level systems
Preferred qualifications
  • PhD in mathematics, physics, computer science, engineering, or a related quantitative field
  • Strong background in linear algebra, probability, numerical methods, or scientific computing
  • Experience working with quantitative research teams or financial models
  • Demonstrated ability to improve real-world inference performance beyond baseline framework or library implementations
  • Familiarity with PTX-level behavior, tensor‑core utilization, or architecture-specific tuning
  • Exposure to ONNX Runtime, TensorRT, Triton, TVM, or similar systems
  • Exposure to neural networks, tree-based models (e.g., LightGBM), state‑space models (e.g., Mamba architectures), and experience with kernel fusion, custom operators, model compilation, or graph-level optimization
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Trading Interview • Northern (KY)

Hybrid
USD 140,000 - 210,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
GPU Performance Architect: Low-Latency CUDA Mastery
GPU Performance Architect: Low-Latency CUDA Mastery

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Inference Optimization Intern – Performance Modeling
Inference Optimization Intern – Performance Modeling

Ifm Us • Sunnyvale (CA)

On-site
USD 30,000 - 60,000
CUDA Kernel Architect for Low-Latency GPU Inference
CUDA Kernel Architect for Low-Latency GPU Inference

Trading Interview • Northern (KY)

Hybrid
USD 140,000 - 210,000