CUDA Engineer - Kernel Optimization - AI Trainer

Mercor

Chicago (IL)

On-site

USD 83,000 - 165,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity targets freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance through profiler-guided analysis.

You will analyze, optimize, and reason about GPU kernels across modern hardware environments, writing C++17 and Python code while applying CUDA, HIP, or related kernel programming techniques to drive

Qualifications

  • Available to work at least 20 hrs/wk.
  • Fluent in core C++ features through C++17.
  • Working knowledge of Python and Git.
  • Fluent in at least one GPU programming model such as CUDA or HIP.
  • At least 1 year of GPU-related experience.
  • Strong understanding of GPU profiler metrics for optimization.
  • Ability to optimize kernels without deep prior context.
  • Experience with CUDA, HIP, inline PTX, or tensor core optimization is a plus.
  • Experience optimizing kernels for NVIDIA Blackwell hardware is a plus.
  • Familiarity with NSight Compute is a plus.
  • Prior GPU hardware experience with NVIDIA/AMD/Qualcomm is a plus.

Responsibilities

  • Analyze and optimize GPU kernels for performance and hardware utilization.
  • Use profiler metrics to guide kernel improvements.
  • Review GPU kernel implementations and identify bottlenecks.
  • Write and modify C++17, Python, and GPU programming code.
  • Apply CUDA, HIP, or related kernel programming to improve performance.
  • Document optimization decisions clearly.

Skills

C++17
Python
GPU programming
CUDA
HIP
NSight Compute
Git
Profiler metrics

Tools

CUDA toolkit
Inline PTX
Tensor cores

Job description

1. Role Overview

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You’ll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures.

2. Key Responsibilities
  • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization
  • Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements
  • Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms
  • Write, modify, and reason about C++17, Python, and GPU programming code
  • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes
  • Document optimization decisions clearly, including when specific profiler metrics are or are not useful
3. Ideal Qualifications
  • Available to work at least 20 hrs/wk
  • Fluent in core C++ features through C++17
  • Working knowledge of Python and Git
  • Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming
  • At least 1 year of professional or graduate-level research experience working with GPUs
  • Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels
  • Ability to optimize GPU kernels without needing deep prior context on every algorithm
  • Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus
  • Experience optimizing kernels for NVIDIA Blackwell hardware is a plus
  • Familiarity with NSight Compute is a plus
  • Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus
  • Open-source contributions related to GPU kernel optimization are a plus
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
CUDA GPU Kernel Performance Engineer (Contract)
CUDA GPU Kernel Performance Engineer (Contract)

Mercor • United States

Remote
GPU Kernel Optimizer for AI Performance (CUDA/C++)
GPU Kernel Optimizer for AI Performance (CUDA/C++)

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Kernel Performance Engineer
CUDA Kernel Performance Engineer

Obsidian • San Francisco (CA)

On-site
USD 80,000 - 120,000
Freelance GPU Kernel Optimizer - CUDA/C++ Specialist
Freelance GPU Kernel Optimizer - CUDA/C++ Specialist

Obsidian • San Francisco (CA)

Remote
USD 90,000 - 130,000
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
KERNEL ENGINEER
KERNEL ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Software Engineer (HPC / Deep Learning Optimization)
GPU Software Engineer (HPC / Deep Learning Optimization)

Luxoft • Town of Poland (NY)

On-site
USD 120,000 - 180,000
Remote CUDA Engineer: GPU Kernel Optimizer & HPC Expert
Remote CUDA Engineer: GPU Kernel Optimizer & HPC Expert

YO HR Consultancy • United States

Remote
USD 120,000 - 180,000