GPU Performance Engineer

Two Sigma

New York (NY)

On-site

USD 120,000 - 190,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Two Sigma seeks a GPU-focused engineer in New York to design and optimize CUDA kernels for financial workloads and to advance precision management across next-generation hardware. You will profile with Nsight and build reusable GPU libraries for modeling teams.

The role requires strong C++ and Python, expert CUDA, and a track record of delivering speedups on real workloads, with experience in multi-GPU and HPC environments.

Qualifications

  • BS or MS in Science, Technology, Engineering or Math.
  • Minimum 1 year of experience required; 4–10 years preferred.
  • Expert-level CUDA programming: kernel development, memory management, stream and graph optimization.
  • Deep understanding of GPU architecture: SM structure, warp scheduling, memory hierarchy (registers, shared memory, L1/L2, HBM).
  • Experience with performance profiling and optimization of GPU workloads.
  • Strong C++ and Python skills, as well as familiarity with mixed-precision computation and numerical stability.
  • Track record of delivering meaningful speedups on real workloads (not just benchmarks).

Responsibilities

  • Design and implement GPU-accelerated kernels for financial computation workloads.
  • Optimize GPU code for throughput, latency, and memory efficiency across current and next-generation hardware (Blackwell, Rubin).
  • Develop procedures for precision management (FP8/FP4 training and inference) in financial applications.
  • Profile and optimize GPU workloads using NVIDIA tooling (Nsight Systems, Nsight Compute).
  • Build reusable GPU libraries and abstractions that modeling teams can use without requiring deep CUDA expertise.
  • Evaluate and integrate GPU-accelerated libraries (RAPIDS, CUTLASS, cuBLAS, TensorRT) for financial use cases.

Skills

CUDA programming
Kernel development
Memory management
Performance profiling
Python
C++
Numerical stability
Multi-GPU programming

Education

Bachelor's or Master's in STEM

Tools

Nsight Systems
Nsight Compute
RAPIDS
CUTLASS
cuBLAS
TensorRT
NCCL
MPI
RAPIDS cuDF

Job description

You will take on the following responsibilities:

  • Design and implement GPU-accelerated kernels for financial computation workloads
  • Optimize GPU code for throughput, latency, and memory efficiency across current and next-generation hardware (Blackwell, Rubin)
  • Develop procedures for precision management (FP8/FP4 training and inference) in financial applications
  • Profile and optimize GPU workloads using NVIDIA tooling (Nsight Systems, Nsight Compute)
  • Build reusable GPU libraries and abstractions that modeling teams can use without requiring deep CUDA expertise
  • Evaluate and integrate GPU-accelerated libraries (RAPIDS, CUTLASS, cuBLAS, TensorRT) for financial use cases

You should possess the following qualifications:

  • BS or MS in Science, Technology, Engineering or Math
  • Minimum 1 year of experience required; 4-10 years of experience preferred
  • Expert-level CUDA programming: kernel development, memory management, stream and graph optimization
  • Deep understanding of GPU architecture: SM structure, warp scheduling, memory hierarchy (registers, shared memory, L1/L2, HBM)
  • Experience with performance profiling and optimization of GPU workloads
  • Strong C++ and Python skills, as well as familiarity with mixed-precision computation and numerical stability
  • Track record of delivering meaningful speedups on real workloads (not just benchmarks)

Preferred experience:

  • Background in HPC, scientific computing, or computational finance
  • Experience with multi-GPU and multi-node GPU programming (NCCL, MPI)
  • Familiarity with GPU-accelerated data processing frameworks (RAPIDS, cuDF)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer — Financial Compute Acceleration
GPU Performance Engineer — Financial Compute Acceleration

Two Sigma • New York (NY)

On-site
USD 120,000 - 190,000
Staff C++ Engineer
Staff C++ Engineer

Glocomms • San Francisco (CA)

On-site
USD 190,000 - 270,000
Software Engineer - GPU Kernel
Software Engineer - GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Senior GPU HPC Engineer - CUDA Optimizations & Equity
Senior GPU HPC Engineer - CUDA Optimizations & Equity

NVIDIA • Arizona

On-site
USD 184,000 - 356,500
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

On-site
USD 200,000 - 300,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Obsidian • San Francisco (CA)

On-site
USD 165,000 - 276,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
Senior Performance Architect - Heterogeneous Workload Optimization
Senior Performance Architect - Heterogeneous Workload Optimization

NVIDIA • Durham (NC)

On-site
USD 184,000 - 287,500
Equity
Comprehensive benefits package