GPU Kernel Optimizer for AI Performance (CUDA/C++)

Mercor

Chicago (IL)

On-site

USD 83,000 - 165,000

Part time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity targets freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance through profiler-guided analysis.

You will analyze, optimize, and reason about GPU kernels across modern hardware environments, writing C++17 and Python code while applying CUDA, HIP, or related kernel programming techniques to drive

Qualifications

  • Available to work at least 20 hrs/wk.
  • Fluent in core C++ features through C++17.
  • Working knowledge of Python and Git.
  • Fluent in at least one GPU programming model such as CUDA or HIP.
  • At least 1 year of GPU-related experience.
  • Strong understanding of GPU profiler metrics for optimization.
  • Ability to optimize kernels without deep prior context.
  • Experience with CUDA, HIP, inline PTX, or tensor core optimization is a plus.
  • Experience optimizing kernels for NVIDIA Blackwell hardware is a plus.
  • Familiarity with NSight Compute is a plus.
  • Prior GPU hardware experience with NVIDIA/AMD/Qualcomm is a plus.

Responsibilities

  • Analyze and optimize GPU kernels for performance and hardware utilization.
  • Use profiler metrics to guide kernel improvements.
  • Review GPU kernel implementations and identify bottlenecks.
  • Write and modify C++17, Python, and GPU programming code.
  • Apply CUDA, HIP, or related kernel programming to improve performance.
  • Document optimization decisions clearly.

Skills

C++17
Python
GPU programming
CUDA
HIP
NSight Compute
Git
Profiler metrics

Tools

CUDA toolkit
Inline PTX
Tensor cores

Job description

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity targets freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance through profiler-guided analysis.

You will analyze, optimize, and reason about GPU kernels across modern hardware environments, writing C++17 and Python code while applying CUDA, HIP, or related kernel programming techniques to drive

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Kernel Optimization Engineer for AI Training
GPU Kernel Optimization Engineer for AI Training

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
GPU Kernel Performance Engineer (Contract)
GPU Kernel Performance Engineer (Contract)

Obsidian • San Francisco (CA)

Remote
USD 83,000 - 138,000
Contract GPU Kernel Optimizer (CUDA/C++17)
Contract GPU Kernel Optimizer (CUDA/C++17)

Obsidian • San Francisco (CA)

On-site
USD 165,000 - 276,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Obsidian • San Francisco (CA)

On-site
USD 165,000 - 276,000
Remote CUDA Kernel Performance Engineer
Remote CUDA Kernel Performance Engineer

Embedded Shishya • United States

Remote
USD 110,000 - 165,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Obsidian • San Francisco (CA)

On-site
USD 80,000 - 120,000