CUDA Engineering Expert

Weekday AI

United States

Remote

USD 110,000 - 138,000

Part time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Weekday AI is seeking a GPU kernel optimization expert for a freelance project with a leading AI lab. You will analyze and optimize GPU kernels, using CUDA/HIP and profilers to squeeze performance across modern hardware.

This contract-based role requires strong C++17, Python, and GPU programming skills and the ability to work remotely on your own schedule. Ideal candidates will have at least 1 year of GPU experience, comfort with Nsight Compute, and a track record of performance improvements.

Qualifications

  • Proficient in modern C++ (C++17)
  • Experience with GPU kernels and performance profiling
  • Familiarity with CUDA/HIP or other GPU programming models
  • Knowledge of GPU hardware architecture and perf metrics

Responsibilities

  • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization
  • Use profiler metrics to guide kernel improvements
  • Review GPU kernel implementations and identify bottlenecks
  • Write, modify, and reason about C++17, Python, and GPU programming code
  • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes
  • Document optimization decisions clearly

Skills

C++17
Python
GPU programming
CUDA
HIP
Shader programming
Profiling / performance analysis

Tools

Git
NSight Compute
CUDA C++ Core Libraries

Job description

This role is for one of our clients


Compensation: $80-$100 per hour


We are seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You’ll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures.


Key Responsibilities


  • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization

  • Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements

  • Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms

  • Write, modify, and reason about C++17, Python, and GPU programming code

  • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes

  • Document optimization decisions clearly, including when specific profiler metrics are or are not useful


Ideal Qualifications


  • Available to work at least 20 hrs/wk

  • Fluent in core C++ features through C++17

  • Working knowledge of Python and Git

  • Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming

  • At least 1 year of professional or graduate-level research experience working with GPUs

  • Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels

  • Ability to optimize GPU kernels without needing deep prior context on every algorithm

  • Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus

  • Experience optimizing kernels for NVIDIA Blackwell hardware is a plus

  • Familiarity with NSight Compute is a plus

  • Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus

  • Open-source contributions related to GPU kernel optimization are a plus


We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.


Contract and Payment Terms


  • You will be engaged as an independent contractor.

  • This is a fully remote role that can be completed on your own schedule.

  • Projects can be extended, shortened, or concluded early depending on needs and performance.

  • Your work will not involve access to confidential or proprietary information from any employer, client, or institution.

  • Payments are weekly on Stripe or Wise based on services rendered.

  • Please note: We are unable to support H1-B or STEM OPT candidates at this time.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

CUDA Engineering Expert
CUDA Engineering Expert

Weekday 1 • United States

Remote
USD 110,000 - 138,000
GPU Programming Expert - Fully Remote | Upto $120/hr
GPU Programming Expert - Fully Remote | Upto $120/hr

mercor • United States

Remote
USD 110,000 - 165,000
CUDA Engineering Expert | Remote - Contract
CUDA Engineering Expert | Remote - Contract

Xperteez Technology Pvt Ltd • United States

Remote
USD 138,000 - 248,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 165,000 - 276,000
GPU Programming Expert - Fully Remote | Upto $500/task Task based
GPU Programming Expert - Fully Remote | Upto $500/task Task based

mercor • United States

Remote
USD 568,000 - 809,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
Remote GPU Kernel Optimization Specialist
Remote GPU Kernel Optimization Specialist

Weekday AI • United States

Remote
USD 110,000 - 138,000
GPU Kernel Performance Engineer (Contract)
GPU Kernel Performance Engineer (Contract)

Obsidian • San Francisco (CA)

Remote
USD 83,000 - 138,000
Remote CUDA Engineer — GPU Kernel Optimizer
Remote CUDA Engineer — GPU Kernel Optimizer

mercor • United States

Remote
USD 110,000 - 165,000
CUDA Software Engineer - Remote
CUDA Software Engineer - Remote

YO IT Consulting • California (MO)

Remote
USD 120,000 - 180,000