GPU Performance Architect: Low-Latency CUDA Mastery

SIG Susquehanna

Pennsylvania

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SIG Susquehanna is seeking a GPU Performance Engineer to optimize CUDA kernels for low-latency inference workloads. You will collaborate with quantitative researchers to enhance performance through low-level optimizations while managing complex model structures.

This position demands a PhD in a quantitative field, solid programming skills in C/C++, and a deep understanding of GPU architectures. Join us in shaping optimized compute pipelines that significantly enhance inference performance.

Qualifications

  • Strong proficiency in writing and optimizing CUDA kernels.
  • Solid programming experience in C/C++.
  • Deep understanding of GPU architecture.
  • Strong problem-solving skills.

Responsibilities

  • Design and optimize custom CUDA kernels.
  • Analyze models to identify computational bottlenecks.
  • Collaborate with researchers for high-performance pipelines.
  • Profile and benchmark GPU performance.
  • Contribute to GPU architecture decisions.

Skills

CUDA optimization
C/C++ programming
GPU architecture understanding
Performance analysis
Numerical stability reasoning

Education

PhD in a quantitative field

Tools

ONNX Runtime
TensorRT
Triton

Job description

SIG Susquehanna is seeking a GPU Performance Engineer to optimize CUDA kernels for low-latency inference workloads. You will collaborate with quantitative researchers to enhance performance through low-level optimizations while managing complex model structures.

This position demands a PhD in a quantitative field, solid programming skills in C/C++, and a deep understanding of GPU architectures. Join us in shaping optimized compute pipelines that significantly enhance inference performance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Kernel Architect for Low-Latency GPU Inference
CUDA Kernel Architect for Low-Latency GPU Inference

Trading Interview • Northern (KY)

Hybrid
USD 140,000 - 210,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
Latency-Critical CUDA Kernel Engineer
Latency-Critical CUDA Kernel Engineer

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
Low-Latency GPU Kernel Engineer
Low-Latency GPU Kernel Engineer

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Trading Interview • Northern (KY)

Hybrid
USD 140,000 - 210,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
ML Inference Engineer - Low-Latency GPU Systems
ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000