CUDA Engineer - Kernel Optimization

Obsidian

San Francisco (CA)

On-site

USD 165,000 - 276,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor seeks GPU kernel optimization experts for a freelance project in the San Francisco area. You will analyze, optimize, and reason about GPU kernels across modern hardware, squeezing performance with profiler-guided analysis.

Requirements include strong C++17, Python, and GPU programming experience (CUDA/HIP). Familiarity with NSight Compute and kernel profiling is a plus; work at least 20 hrs/week on contract basis.

Qualifications

  • Fluent in core C++ features through C++17.
  • Working knowledge of Python and Git.
  • Fluent in at least one GPU programming model (CUDA, HIP, Slang, HLSL, GLSL).
  • At least 1 year of GPU-related research or professional experience.
  • Understanding of GPU profiler metrics and how to optimize kernels.

Responsibilities

  • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization.
  • Use profiler metrics (e.g., L2 cache hit rate, L2 throughput, occupancy) to guide kernel improvements.
  • Review kernel implementations and identify bottlenecks with minimal context on algorithms.
  • Write, modify, and reason about C++17, Python, and GPU programming code.
  • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve outcomes.
  • Document optimization decisions clearly, including usefulness of specific profiler metrics.

Skills

C++17
Python
Git
GPU profiling
CUDA

Tools

NSight Compute

Job description

1. Role Overview

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You'll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures.

2. Key Responsibilities
  • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization

  • Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements

  • Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms

  • Write, modify, and reason about C++17, Python, and GPU programming code

  • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes

  • Document optimization decisions clearly, including when specific profiler metrics are or are not useful

3. Ideal Qualifications
  • Available to work at least 20 hrs/wk

  • Fluent in core C++ features through C++17

  • Working knowledge of Python and Git

  • Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming

  • At least 1 year of professional or graduate-level research experience working with GPUs

  • Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels

  • Ability to optimize GPU kernels without needing deep prior context on every algorithm

  • Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus

  • Experience optimizing kernels for NVIDIA Blackwell hardware is a plus

  • Familiarity with NSight Compute is a plus

  • Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus

  • Open-source contributions related to GPU kernel optimization are a plus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
CUDA GPU Kernel Performance Engineer (Contract)
CUDA GPU Kernel Performance Engineer (Contract)

Mercor • United States

Remote
GPU Kernel Performance Engineer (Contract)
GPU Kernel Performance Engineer (Contract)

Obsidian • San Francisco (CA)

Remote
USD 83,000 - 138,000
GPU Kernel Optimization Engineer for AI Training
GPU Kernel Optimization Engineer for AI Training

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
GPU Kernel Optimizer for AI Performance (CUDA/C++)
GPU Kernel Optimizer for AI Performance (CUDA/C++)

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Obsidian • San Francisco (CA)

On-site
USD 80,000 - 120,000
Contract GPU Kernel Optimizer (CUDA/C++17)
Contract GPU Kernel Optimizer (CUDA/C++17)

Obsidian • San Francisco (CA)

On-site
USD 165,000 - 276,000