AI Kernel Quality Engineer — GPU/Accelerator Expert

Obsidian

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian is seeking a skilled reviewer to evaluate GPU/accelerator kernel development tasks used to train frontier AI models. You will assess numerical correctness, performance benchmarks, and task scoping, ensuring compilation and runtime validity across diverse kernels.

The role requires 3+ years in CUDA, Triton, NKI, or Pallas, plus experience with profiling tools and handling common failure modes. This is a research-oriented, full-time position in San Francisco.

Qualifications

  • 3+ years of hands-on experience developing, optimizing, or verifying GPU/accelerator kernels in two or more of CUDA, Triton, NKI or Pallas (JAX).
  • Strong understanding of numerical-correctness criteria for kernels (absolute/relative/ULP tolerances, reference-improvement selection).
  • Experience with performance profiling and benchmarking (Nsight, ncu, roofline analysis or framework profilers).
  • Familiarity with common compilation and runtime failure modes (driver mismatches, OOM, launch-configuration errors, shape/stride mismatches).
  • Experience with at least three kernel task types: generation from specification, translation/lowering, migration across hardware targets, debugging, optimization, or operator fusion

Responsibilities

  • Evaluate the quality, correctness, and completeness of GPU/accelerator kernel development tasks.
  • Provide rubric-based written feedback to train and assess a frontier AI lab's models.
  • Assess compilation/runtime validity across diverse kernel task types.
  • Produce actionable recommendations to improve task scoping and benchmarking fairness

Skills

CUDA
Triton
NKI
Pallas (JAX)

Tools

Nsight
ncu
Roofline

Job description

Obsidian is seeking a skilled reviewer to evaluate GPU/accelerator kernel development tasks used to train frontier AI models. You will assess numerical correctness, performance benchmarks, and task scoping, ensuring compilation and runtime validity across diverse kernels.

The role requires 3+ years in CUDA, Triton, NKI, or Pallas, plus experience with profiling tools and handling common failure modes. This is a research-oriented, full-time position in San Francisco.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Kernel Quality Engineer
GPU Kernel Quality Engineer

Obsidian • New York (NY)

Remote
USD 140,000 - 200,000
AI GPU Kernel Engineer: Performance & Validation
AI GPU Kernel Engineer: Performance & Validation

Obsidian • Chicago (IL)

On-site
USD 120,000 - 180,000
GPU Kernel Architect for AI Training & Evaluation
GPU Kernel Architect for AI Training & Evaluation

Mercor • Chicago (IL)

On-site
USD 120,000 - 180,000
GPU Kernel Quality & Performance Engineer
GPU Kernel Quality & Performance Engineer

Mercor • New York (NY)

Remote
USD 150,000 - 190,000
GPU Kernel Expert - AI Specialist
GPU Kernel Expert - AI Specialist

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
GPU Kernel Expert - AI Specialist
GPU Kernel Expert - AI Specialist

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000
GPU Kernel Developer - AI Trainer
GPU Kernel Developer - AI Trainer

Mercor • Chicago (IL)

On-site
USD 120,000 - 180,000
GPU Kernel Developer - AI Trainer
GPU Kernel Developer - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 120,000 - 180,000
Lead GPU Kernel & AI Performance Engineer
Lead GPU Kernel & AI Performance Engineer

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
GPU Kernel Evaluation Expert
GPU Kernel Evaluation Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000