Remote GPU Kernel Engineer: CUDA/Triton Optimization

anyone-ai

United States

On-site

USD 138,000 - 248,000

Part time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Anyone AI is recruiting experienced GPU Kernel Engineers for a remote, part-time, project-based consulting engagement focused on reviewing, debugging, and evaluating AI compute kernels.

You’ll work on CUDA/Triton kernels, assess numerical correctness, performance, and memory usage, and provide actionable feedback to optimize for target hardware, including translation between frameworks and hardware migrations, with emphasis on reproducibility and benchmarking.

Qualifications

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU/accelerator kernels.
  • Strong understanding of GPU performance optimization.
  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers.
  • Understanding of floating-point numerical correctness and tolerance thresholds.
  • Experience debugging kernel compilation and runtime issues.
  • Ability to distinguish software defects, environment problems, and genuine optimization challenges.

Responsibilities

  • Reviewing GPU and accelerator kernel implementations for correctness.
  • Comparing outputs against reference implementations.
  • Evaluating numerical tolerance thresholds.
  • Reviewing kernel benchmarks and determining whether comparisons are fair.
  • Identifying performance bottlenecks and optimization opportunities.
  • Assessing whether performance targets are realistic given hardware limits.
  • Reviewing kernel translations and hardware migrations.
  • Identifying compilation, memory, and runtime issues.
  • Providing clear, actionable technical feedback.

Skills

GPU kernel development
Kernel optimization
Kernel debugging
GPU performance analysis
Profiling tools experience
Numerical correctness

Tools

Nsight
NCU
Roofline analysis
Framework profilers
Triton
Pallas

Job description

Anyone AI is recruiting experienced GPU Kernel Engineers for a remote, part-time, project-based consulting engagement focused on reviewing, debugging, and evaluating AI compute kernels.

You’ll work on CUDA/Triton kernels, assess numerical correctness, performance, and memory usage, and provide actionable feedback to optimize for target hardware, including translation between frameworks and hardware migrations, with emphasis on reproducibility and benchmarking.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance
GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

anyone-ai • United States

Remote
USD 138,000 - 248,000
Remote CUDA Kernel Engineer – Optimize GPU Performance
Remote CUDA Kernel Engineer – Optimize GPU Performance

Pragmatike • Cambridge (MA)

On-site
USD 150,000 - 230,000
Salary + equity
Sign-on bonus
Health/Dental/Vision
+1
Remote GPU Kernel Evaluator & Optimizer
Remote GPU Kernel Evaluator & Optimizer

Appsierra Group • United States

On-site
USD 146,000 - 187,000
Fully remote
Flexible schedule
Weekly payments via Stripe or Wise
Remote CUDA Kernel & GPU Performance Expert
Remote CUDA Kernel & GPU Performance Expert

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 83,000 - 138,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
GPU Kernel Architect for AI Training & Evaluation
GPU Kernel Architect for AI Training & Evaluation

Mercor • Chicago (IL)

On-site
USD 120,000 - 180,000
Remote | CUDA Engineering Expert — $60–$100/hour
Remote | CUDA Engineering Expert — $60–$100/hour

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 83,000 - 138,000
GPU Systems Research Intern (CUDA/Triton)
GPU Systems Research Intern (CUDA/Triton)

Togetherai • San Francisco (CA)

On-site
USD 80,000 - 96,000
Housing stipend
Competitive compensation
Remote CUDA Kernel Performance Engineer
Remote CUDA Kernel Performance Engineer

Embedded Shishya • United States

Remote
USD 110,000 - 165,000