GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

anyone-ai

United States

Remote

USD 138,000 - 248,000

Part time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Anyone AI is recruiting experienced GPU Kernel Engineers for a remote, part-time, project-based consulting engagement focused on reviewing, debugging, and evaluating AI compute kernels.

You’ll work on CUDA/Triton kernels, assess numerical correctness, performance, and memory usage, and provide actionable feedback to optimize for target hardware, including translation between frameworks and hardware migrations, with emphasis on reproducibility and benchmarking.

Qualifications

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU/accelerator kernels.
  • Strong understanding of GPU performance optimization.
  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers.
  • Understanding of floating-point numerical correctness and tolerance thresholds.
  • Experience debugging kernel compilation and runtime issues.
  • Ability to distinguish software defects, environment problems, and genuine optimization challenges.

Responsibilities

  • Reviewing GPU and accelerator kernel implementations for correctness.
  • Comparing outputs against reference implementations.
  • Evaluating numerical tolerance thresholds.
  • Reviewing kernel benchmarks and determining whether comparisons are fair.
  • Identifying performance bottlenecks and optimization opportunities.
  • Assessing whether performance targets are realistic given hardware limits.
  • Reviewing kernel translations and hardware migrations.
  • Identifying compilation, memory, and runtime issues.
  • Providing clear, actionable technical feedback.

Skills

GPU kernel development
Kernel optimization
Kernel debugging
GPU performance analysis
Profiling tools experience
Numerical correctness

Tools

Nsight
NCU
Roofline analysis
Framework profilers
Triton
Pallas

Job description

Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas , with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

  • Kernel implementation and debugging

  • CUDA and Triton optimization

  • Translation between kernel frameworks

  • Hardware migration

  • Operator fusion

  • Performance profiling and benchmarking

  • Numerical correctness verification

  • Compilation and runtime debugging

  • Memory hierarchy optimization

  • Kernel-level AI workload performance

You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

What We’re Looking For
  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels

  • Strong experience with at least two of the following:

  • Strong understanding of GPU performance optimization

  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers

  • Understanding of:

  • Strong understanding of floating-point numerical correctness and tolerance thresholds

  • Experience debugging kernel compilation and runtime issues

  • Ability to distinguish software defects, environment problems, and genuine optimization challenges

Relevant Experience

Candidates should have experience with several of the following types of work:

  • Writing kernels from technical specifications

  • Translating kernels between CUDA, Triton, or other frameworks

  • Migrating kernels across hardware platforms

  • Debugging incorrect kernel implementations

  • Optimizing kernel performance

  • Fusing multiple operations into optimized kernels

Nice to Have
  • Experience across both NVIDIA GPU and custom accelerator ecosystems

  • Experience with AWS Trainium, TPU, JAX, or other accelerators

  • Compiler engineering experience

  • Familiarity with MLIR, XLA, or intermediate representation lowering

  • Contributions to GPU or ML kernel libraries

  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls

  • Experience with AI model evaluation, RLHF, or technical benchmark development

What You’ll Be Responsible For
  • Reviewing GPU and accelerator kernel implementations for correctness

  • Comparing outputs against reference implementations

  • Evaluating numerical tolerance thresholds

  • Reviewing kernel benchmarks and determining whether comparisons are fair

  • Identifying performance bottlenecks and optimization opportunities

  • Assessing whether performance targets are realistic given hardware limits

  • Reviewing kernel translations and hardware migrations

  • Identifying compilation, driver, memory, shape, and runtime issues

  • Determining whether technical tasks are genuinely difficult or incorrectly configured

  • Providing clear, actionable technical feedback

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote GPU Kernel Engineer: CUDA/Triton Optimization
Remote GPU Kernel Engineer: CUDA/Triton Optimization

anyone-ai • United States

On-site
USD 138,000 - 248,000
Remote | CUDA Engineering Expert — $60–$100/hour
Remote | CUDA Engineering Expert — $60–$100/hour

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 83,000 - 138,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Kernel Engineer: 26-02794
Senior Kernel Engineer: 26-02794

Akraya, Inc. • Bellevue (WA)

On-site
USD 117,000 - 124,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Obsidian • San Francisco (CA)

On-site
USD 165,000 - 276,000
GPU Kernel Expert - AI Specialist
GPU Kernel Expert - AI Specialist

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000
GPU Kernel Expert - AI Specialist
GPU Kernel Expert - AI Specialist

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000