CUDA Engineer - Kernel Optimization

Mercor

Mumbai

On-site

INR 1,653,000 - 3,306,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based role requires strong C++ skills, practical GPU programming experience, and the ability to squeeze performance out of modern GPU architectures.

You will analyze, optimize, and reason about GPU kernels across contemporary hardware, using profiler-driven metrics to guide improvements. Familiarity with CUDA/HIP and Python is valued.

Qualifications

  • Available to work at least 20 hrs/wk.
  • Fluent in core C++ features through C++17.
  • Working knowledge of Python and Git.
  • Fluent in at least one GPU programming model such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming.
  • At least 1 year of professional or graduate-level research experience working with GPUs.
  • Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels.
  • Ability to optimize GPU kernels without needing deep prior context on every algorithm.
  • Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus.
  • Experience optimizing kernels for NVIDIA Blackwell hardware is a plus.
  • Familiarity with NSight Compute is a plus.
  • Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus.
  • Open-source contributions related to GPU kernel optimization are a plus.

Responsibilities

  • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization
  • Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements
  • Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms
  • Write, modify, and reason about C++17, Python, and GPU programming code
  • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes
  • Document optimization decisions clearly, including when specific profiler metrics are or are not useful

Skills

C++17 features
Python
GPU kernel optimization
Profiler-guided analysis
CUDA
HIP
Shader programming
Kernel performance reasoning
Tensor core optimization

Tools

Git
NSight Compute
CUDA toolkit
Slang
HLSL
GLSL

Job description

1. Role Overview

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You'll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures.

2. Key Responsibilities
  • Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization

  • Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements

  • Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms

  • Write, modify, and reason about C++17, Python, and GPU programming code

  • Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes

  • Document optimization decisions clearly, including when specific profiler metrics are or are not useful

3. Ideal Qualifications
  • Available to work at least 20 hrs/wk

  • Fluent in core C++ features through C++17

  • Working knowledge of Python and Git

  • Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming

  • At least 1 year of professional or graduate-level research experience working with GPUs

  • Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels

  • Ability to optimize GPU kernels without needing deep prior context on every algorithm

  • Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus

  • Experience optimizing kernels for NVIDIA Blackwell hardware is a plus

  • Familiarity with NSight Compute is a plus

  • Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus

  • Open-source contributions related to GPU kernel optimization are a plus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Obsidian • Mumbai

On-site
INR 2,066,000 - 4,133,000
GPU Programming Expert - Fully Remote
GPU Programming Expert - Fully Remote

Mercor • Mumbai

Remote
INR 1,653,000 - 3,444,000
GPU Programming Expert - Fully Remote | Upto $500/task Task based
GPU Programming Expert - Fully Remote | Upto $500/task Task based

Obsidian • Mumbai

Remote
INR 2,755,000 - 4,822,000
Hardware Engineer (Remote | $80 –$100/hr)
Hardware Engineer (Remote | $80 –$100/hr)

Synthires • India

On-site
INR 10,526,000 - 13,158,000
CUDA Kernel Engineer
CUDA Kernel Engineer

Nava • Bengaluru

On-site
INR 1,500,000 - 2,600,000
Solution Architect - GPU/TPU Kernel Optimization
Solution Architect - GPU/TPU Kernel Optimization

EPAM Systems • Hyderabad

On-site
INR 2,500,000 - 3,500,000
Solution Architect - GPU/TPU Kernel Optimization
Solution Architect - GPU/TPU Kernel Optimization

EPAM Systems • Chennai District

On-site
INR 2,000,000 - 3,000,000
GPU Optimization Engineer
GPU Optimization Engineer

Xevyte Technologies • Bengaluru

Hybrid
INR 1,500,000 - 2,100,000
GPU Optimisation Engineer
GPU Optimisation Engineer

Xopuntech (india) • Bengaluru

On-site
INR 2,400,000 - 4,800,000
Performance Engineer, Kernels
Performance Engineer, Kernels

Sarvam • Bengaluru

Hybrid
INR 4,000,000 - 8,000,000