GPU Performance Engineer: CUDA Kernel Optimizer

Two Sigma

New York (NY)

Hybrid

USD 165,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Dental insurance
401k match
Life & disability insurance
Hybrid work policy
Tuition reimbursement
Conference sponsorship
Vacation & leave
Wellness activities
Onsite gyms

Job summary

Two Sigma seeks a GPU programming expert to lead the design of GPU-accelerated kernels for financial workloads. You will optimize culture-changing code, manage precision (FP8/FP4), and profile performance with NVIDIA tools across current and next-gen hardware.

You will build reusable GPU libraries and evaluate RAPIDS, CUTLASS, cuBLAS, TensorRT, NCCL, MPI for real-world financial use cases. A strong background in C++/Python and GPU architecture is essential.

Qualifications

  • BS or MS in Science, Technology, Engineering or Math.
  • Minimum 1 year of experience; 4–10 years preferred.
  • Expert-level CUDA programming: kernel development, memory management, stream and graph optimization.
  • Deep understanding of GPU architecture: SM structure, warp scheduling, memory hierarchy.
  • Experience with performance profiling and optimization of GPU workloads.
  • Strong C++ and Python skills, plus familiarity with mixed-precision computation and numerical stability.
  • Track record of delivering speedups on real workloads (not just benchmarks).

Responsibilities

  • Design and implement GPU-accelerated kernels for financial computation workloads
  • Optimize GPU code for throughput, latency, and memory efficiency across current and next-generation hardware (Blackwell, Rubin)
  • Develop procedures for precision management (FP8/FP4 training and inference) in financial applications
  • Profile and optimize GPU workloads using NVIDIA tooling (Nsight Systems, Nsight Compute)
  • Build reusable GPU libraries and abstractions that modeling teams can use without requiring deep CUDA expertise
  • Evaluate and integrate GPU-accelerated libraries (RAPIDS, CUTLASS, cuBLAS, TensorRT) for financial use cases

Skills

CUDA programming
C++ & Python
Parallel computing
Performance optimization
GPU architecture
Profiling GPU workloads

Education

BS/MS in science/engineering/math

Tools

Nsight Systems
Nsight Compute
RAPIDS
cuBLAS
TensorRT
CUTLASS
RAPIDS cuDF
NCCL
MPI

Job description

Two Sigma seeks a GPU programming expert to lead the design of GPU-accelerated kernels for financial workloads. You will optimize culture-changing code, manage precision (FP8/FP4), and profile performance with NVIDIA tools across current and next-gen hardware.

You will build reusable GPU libraries and evaluate RAPIDS, CUTLASS, cuBLAS, TensorRT, NCCL, MPI for real-world financial use cases. A strong background in C++/Python and GPU architecture is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Performance Engineer
GPU Performance Engineer

Two Sigma • New York (NY)

Hybrid
USD 165,000 - 300,000
Medical insurance
Dental insurance
401k match
+7
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
Freelance GPU Kernel Optimizer - CUDA/C++ Specialist
Freelance GPU Kernel Optimizer - CUDA/C++ Specialist

Obsidian • San Francisco (CA)

Remote
USD 90,000 - 130,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Obsidian • San Francisco (CA)

On-site
USD 80,000 - 120,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
CUDA GPU Kernel Performance Engineer (Contract)
CUDA GPU Kernel Performance Engineer (Contract)

Mercor • United States

Remote
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000