GPU Kernel Developer - AI Trainer

Obsidian

Chicago (IL)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian in Chicago seeks a specialist to evaluate GPU/accelerator kernel development tasks used to train frontier AI models. You will judge numerical correctness, performance benchmarks, task scoping, and compile/run validity across diverse kernel task types.

This role demands rubric-based written feedback and clear guidance to improve kernels, with emphasis on cross-framework translation and hardware-target migrations.

Qualifications

  • 3+ years hands-on experience developing, optimizing, or verifying GPU kernels in at least two of CUDA, Triton, NKI, or Pallas (JAX).
  • Strong understanding of numerical-correctness criteria for kernels (absolute/relative/ULP tolerances, reference-implementation selection).
  • Experience with performance profiling and benchmarking (nsight, ncu, roofline analysis, or framework-native profilers).
  • Familiarity with common compilation/runtime failure modes (driver mismatches, OOM, launch-configuration errors, shape/stride mismatches, autotuning failures).
  • Experience with at least three kernel task types: generation from specification, translation/lowering across frameworks, migration between hardware targets, debugging, performance optimization, or operator fusion.

Skills

CUDA
Triton
NKI
Pallas (JAX)
Performance profiling

Tools

Nsight
NCU

Job description

Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks used to train and evaluate a frontier AI lab's models. You'll assess numerical correctness, performance-benchmarking fairness, task scoping, and compilation/runtime validity across diverse kernel task types — and provide clear, rubric-based written feedback.

Basic Qualifications
  • 3+ years of hands-on experience developing, optimizing, or verifying GPU/accelerator kernels in at least two of: CUDA, Triton, NKI, or Pallas (JAX)
  • Strong understanding of numerical-correctness criteria for kernels (absolute/relative/ULP tolerances, reference-implementation selection)
  • Demonstrated experience with performance profiling and benchmarking (nsight, ncu, roofline analysis, or framework-native profilers)
  • Familiarity with common compilation and runtime failure modes (driver mismatches, OOM, launch-configuration errors, shape/stride mismatches, autotuning failures)
  • Experience with at least three kernel task types: generation from specification, translation/lowering across frameworks, migration between hardware targets, debugging, performance optimization, or operator fusion
Preferred Qualifications
  • Experience across both NVIDIA GPU (CUDA/Triton) and custom-accelerator (NKI/Pallas/TPU) ecosystems
  • Background in compiler engineering, MLIR, or intermediate-representation lowering
  • Understanding of memory-hierarchy optimization (shared-memory tiling, register pressure, bank conflicts, coalescing patterns)
  • Contributions to kernel libraries (cuBLAS, cuDNN, Triton community kernels, JAX/XLA custom calls)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Kernel Developer - AI Trainer
GPU Kernel Developer - AI Trainer

Mercor • Chicago (IL)

On-site
USD 120,000 - 180,000
GPU Kernel Expert - AI Specialist
GPU Kernel Expert - AI Specialist

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
GPU Kernel Expert - AI Specialist
GPU Kernel Expert - AI Specialist

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000
GPU Kernel Expert Mercor · Remote — United States $70-90/hr →
GPU Kernel Expert Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 100,000 - 180,000
GPU Kernel Architect for AI Training & Evaluation
GPU Kernel Architect for AI Training & Evaluation

Mercor • Chicago (IL)

On-site
USD 120,000 - 180,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
GPU Kernel Quality & Performance Engineer
GPU Kernel Quality & Performance Engineer

Mercor • New York (NY)

Remote
USD 150,000 - 190,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
GPU Kernel Evaluation Expert
GPU Kernel Evaluation Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
AI Kernel Quality Engineer — GPU/Accelerator Expert
AI Kernel Quality Engineer — GPU/Accelerator Expert

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000