CUDA Engineer - Kernel Optimization - AI Trainer

Obsidian

Chicago (IL)

On-site

USD 83,000 - 165,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity targets freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis.

You will evaluate, optimize, and reason about GPU kernels across modern hardware environments, writing C++17, Python, and CUDA/HIP code, and documenting optimization decisions clearly.

Qualifications

  • Strong C++ skills and GPU programming experience.
  • Experience with profiler-guided optimization and kernel performance analysis.
  • Familiarity with NVIDIA/AMD/Qualcomm GPU architectures is a plus.

Responsibilities

  • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization.
  • Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy to guide kernel improvements.
  • Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms.
  • Write, modify, and reason about C++17, Python, and GPU programming code.
  • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes.
  • Document optimization decisions clearly, including when specific profiler metrics are or are not useful.

Skills

C++17
GPU programming
Profiling skills
Python

Tools

CUDA
HIP
Slang
HLSL
GLSL
PTX
Tensor cores
NSight Compute
CUDA C++ Core Libraries

Job description

1. Role Overview

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You'll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures.

2. Key Responsibilities
  • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization

  • Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements

  • Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms

  • Write, modify, and reason about C++17, Python, and GPU programming code

  • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes

  • Document optimization decisions clearly, including when specific profiler metrics are or are not useful

3. Ideal Qualifications
  • Available to work at least 20 hrs/wk

  • Fluent in core C++ features through C++17

  • Working knowledge of Python and Git

  • Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming

  • At least 1 year of professional or graduate-level research experience working with GPUs

  • Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels

  • Ability to optimize GPU kernels without needing deep prior context on every algorithm

  • Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus

  • Experience optimizing kernels for NVIDIA Blackwell hardware is a plus

  • Familiarity with NSight Compute is a plus

  • Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus

  • Open-source contributions related to GPU kernel optimization are a plus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Obsidian • San Francisco (CA)

On-site
USD 165,000 - 276,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
GPU Kernel Optimization Engineer for AI Training
GPU Kernel Optimization Engineer for AI Training

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
CUDA GPU Kernel Performance Engineer (Contract)
CUDA GPU Kernel Performance Engineer (Contract)

Mercor • United States

Remote
GPU Kernel Optimizer for AI Performance (CUDA/C++)
GPU Kernel Optimizer for AI Performance (CUDA/C++)

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
GPU Kernel Performance Engineer (Contract)
GPU Kernel Performance Engineer (Contract)

Obsidian • San Francisco (CA)

Remote
USD 83,000 - 138,000
GPU Kernel Expert Mercor · Remote — United States $70-90/hr →
GPU Kernel Expert Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 100,000 - 180,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Obsidian • San Francisco (CA)

On-site
USD 80,000 - 120,000