GPU Kernel Developer - AI Trainer

Mercor

Chicago (IL)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor seeks a professional to evaluate GPU/accelerator kernel development tasks used in frontier AI model training and evaluation. You will assess numerical correctness, performance benchmarking fairness, task scoping, and compilation/runtime validity across diverse kernel types.

The role requires strong experience with CUDA/Triton/NKI/Pallas ecosystems, profilers, and common failure modes. You will deliver rubric-based written feedback to guide improvements.

Qualifications

  • 3+ years developing, optimizing, or verifying GPU/kernel code in two or more frameworks.
  • Strong understanding of numerical-correctness criteria and reference implementations.
  • Experience with performance profiling and benchmarking tools.
  • Familiarity with common kernel compilation/runtime failure modes.
  • Experience with at least three kernel task types: generation, translation/lowering, migration, debugging, optimization, or fusion.

Responsibilities

  • Evaluate kernel development tasks for correctness, performance, and completeness.
  • Provide rubric-based written feedback on scope, correctness, and runtime behavior.

Skills

GPU kernels
CUDA
Triton
NKI/Pallas (JAX)
numerical correctness
performance profiling
benchmarking
kernel task types

Tools

Nsight
ncu
roofline analysis

Job description

Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks used to train and evaluate a frontier AI lab's models. You'll assess numerical correctness, performance-benchmarking fairness, task scoping, and compilation/runtime validity across diverse kernel task types — and provide clear, rubric-based written feedback.

Basic Qualifications
  • 3+ years of hands-on experience developing, optimizing, or verifying GPU/accelerator kernels in at least two of: CUDA, Triton, NKI, or Pallas (JAX)
  • Strong understanding of numerical-correctness criteria for kernels (absolute/relative/ULP tolerances, reference-implementation selection)
  • Demonstrated experience with performance profiling and benchmarking (nsight, ncu, roofline analysis, or framework-native profilers)
  • Familiarity with common compilation and runtime failure modes (driver mismatches, OOM, launch-configuration errors, shape/stride mismatches, autotuning failures)
  • Experience with at least three kernel task types: generation from specification, translation/lowering across frameworks, migration between hardware targets, debugging, performance optimization, or operator fusion
Preferred Qualifications
  • Experience across both NVIDIA GPU (CUDA/Triton) and custom-accelerator (NKI/Pallas/TPU) ecosystems
  • Background in compiler engineering, MLIR, or intermediate-representation lowering
  • Understanding of memory-hierarchy optimization (shared-memory tiling, register pressure, bank conflicts, coalescing patterns)
  • Contributions to kernel libraries (cuBLAS, cuDNN, Triton community kernels, JAX/XLA custom calls)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Kernel Developer - AI Trainer
GPU Kernel Developer - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 120,000 - 180,000
GPU Kernel Expert - AI Specialist
GPU Kernel Expert - AI Specialist

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000
GPU Kernel Expert - AI Specialist
GPU Kernel Expert - AI Specialist

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
GPU Kernel Expert Mercor · Remote — United States $70-90/hr →
GPU Kernel Expert Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 100,000 - 180,000
GPU Kernel Architect for AI Training & Evaluation
GPU Kernel Architect for AI Training & Evaluation

Mercor • Chicago (IL)

On-site
USD 120,000 - 180,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
GPU Kernel Quality & Performance Engineer
GPU Kernel Quality & Performance Engineer

Mercor • New York (NY)

Remote
USD 150,000 - 190,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
GPU Kernel Evaluation Expert
GPU Kernel Evaluation Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
AI Kernel Quality Engineer — GPU/Accelerator Expert
AI Kernel Quality Engineer — GPU/Accelerator Expert

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000