Remote | GPU Kernel Engineer — $60–$80/hour

24-Mag Llc

New York, Northern (NY, KY)

Hybrid

USD 83,000 - 110,000

Part time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

24-MAG LLC offers a specialised part-time consulting opportunity for experienced GPU and accelerator engineers with hands-on expertise in kernel development, numerical validation, performance optimisation, profiling, and low-level compute frameworks. This role focuses on evaluating GPU and accelerator kernel-development tasks for correctness, completeness, reproducibility, and performance quality.

You will review kernel implementations, benchmark methodology, compilation and runtime behaviour,

Qualifications

  • 3+ years hands-on experience developing, optimising, or verifying GPU/accelerator kernels.
  • Experience with at least two of CUDA, Triton, NKI, or Pallas (JAX).
  • Strong numerical correctness concepts including absolute, relative, and ULP tolerances.
  • Proven performance profiling, benchmarking, and debugging skills.

Responsibilities

  • Evaluate GPU and accelerator kernels for technical correctness.
  • Review kernel implementations from specifications or reference operators.
  • Assess numerical behaviour and tolerances.
  • Benchmark methodology and runtime behaviour review.
  • Evaluate memory tiling, coalescing, and efficiency.
  • Identify bottlenecks and ensure numerical correctness during optimisations.

Skills

GPU kernel development
Numerical correctness
Performance profiling
Debugging kernels
Technical feedback

Tools

Nsight
Nsight Compute
roofline analysis
CUDA
Triton
NKI
Pallas (JAX)

Job description

We are sharing a specialised part-time consulting opportunity for experienced GPU and accelerator engineers with hands-on expertise in kernel development, numerical validation, performance optimisation, profiling, and low-level compute frameworks.

This role focuses on evaluating GPU and accelerator kernel-development tasks for correctness, completeness, reproducibility, and performance quality. Selected engineers will review kernel implementations, benchmark methodology, compilation and runtime behaviour, numerical tolerances, and optimisation decisions across multiple accelerator ecosystems.

Key Responsibilities
GPU Kernel Development Review
  • Evaluate GPU and accelerator kernels for technical correctness and completeness
  • Review implementations developed from specifications or reference operators
  • Assess whether kernels correctly implement intended mathematical behaviour
  • Identify implementation defects, unsupported assumptions, or incomplete solutions
  • Apply practical judgement grounded in hands-on kernel engineering experience
CUDA & Accelerator Frameworks
  • Review kernel development across frameworks such as CUDA, Triton, NKI, and Pallas (JAX)
  • Assess framework-specific implementation choices and execution constraints
  • Evaluate kernel translations or migrations between different frameworks
  • Identify incorrect assumptions when moving implementations across accelerator ecosystems
  • Compare alternative kernel implementations for correctness and technical quality
Numerical Correctness
  • Assess kernel outputs against suitable reference implementations
  • Evaluate absolute, relative, and ULP-based tolerances
  • Determine whether numerical differences fall within acceptable limits
  • Review floating-point behaviour and precision-related edge cases
  • Identify discrepancies caused by implementation defects rather than expected numerical variation
Performance Profiling & Benchmarking
  • Evaluate kernel performance using tools such as Nsight, Nsight Compute (ncu), roofline analysis, or framework-native profilers
  • Assess whether benchmark methodology produces fair and meaningful comparisons
  • Review latency, throughput, utilisation, and memory behaviour
  • Identify misleading benchmarking practices or inappropriate baselines
  • Determine whether claimed performance improvements are supported by evidence
Kernel Performance Optimisation
  • Review optimisation strategies for compute and memory efficiency
  • Assess tiling, vectorisation, parallelisation, and workload decomposition
  • Evaluate trade-offs between arithmetic throughput and memory movement
  • Identify bottlenecks affecting kernel performance
  • Assess whether optimisations preserve numerical correctness
Memory Hierarchy Optimisation
  • Review use of registers, shared memory, caches, and accelerator-specific memory resources
  • Evaluate shared-memory tiling, register pressure, bank conflicts, and coalescing patterns
  • Identify inefficient memory-access behaviour
  • Assess data locality and memory-bandwidth utilisation
  • Evaluate whether memory optimisations appropriately match the target hardware
Compilation & Runtime Validation
  • Diagnose common kernel compilation and runtime failures
  • Review issues involving driver incompatibilities, out-of-memory conditions, launch configurations, shape or stride mismatches, and autotuning failures
  • Determine whether failures originate from kernel logic, environment configuration, or runtime assumptions
  • Evaluate proposed debugging approaches and corrective actions
  • Assess whether tasks execute reliably in their intended environment
Kernel Translation & Hardware Migration
  • Review kernels translated or lowered across programming frameworks
  • Evaluate migration between different accelerator targets
  • Assess whether computational semantics and performance assumptions remain valid
  • Identify platform-specific behaviour that requires redesign rather than direct translation
  • Evaluate migration quality across GPU and custom-accelerator environments
Debugging & Operator Fusion
  • Review debugging tasks involving incorrect or unstable kernel implementations
  • Diagnose failures using outputs, profiler data, runtime behaviour, and source code
  • Evaluate operator-fusion strategies where relevant
  • Assess whether fused kernels preserve intended semantics
  • Identify optimisation decisions that introduce correctness or maintainability issues
Compiler & Lowering Concepts
  • Evaluate kernel tasks involving compiler or intermediate-representation concepts where applicable
  • Review transformations between high-level operators and accelerator-level implementations
  • Assess lowering decisions for correctness and efficiency
  • Apply familiarity with MLIR or comparable compiler infrastructures where relevant
  • Identify issues arising from compiler or code-generation assumptions
Rubric-Based Technical Evaluation
  • Assess assigned kernel tasks against structured technical criteria
  • Provide clear written explanations supporting evaluation decisions
  • Reference specific numerical, performance, compilation, or runtime evidence
  • Apply evaluation standards consistently across assignments
  • Distinguish valid implementation alternatives from technically flawed approaches
Ideal Profile
  • 3+ years of hands-on experience developing, optimising, or verifying GPU or accelerator kernels
  • Practical experience with at least two of CUDA, Triton, NKI, or Pallas (JAX)
  • Strong understanding of numerical correctness, including absolute, relative, and ULP tolerances
  • Experience selecting and validating appropriate reference implementations
  • Strong performance profiling and benchmarking experience
  • Familiarity with Nsight, Nsight Compute, roofline analysis, or comparable profiling tools
  • Strong understanding of common kernel compilation and runtime failure modes
  • Experience with at least three of the following:
    • Kernel generation from specification
    • Framework translation or lowering
    • Hardware-target migration
    • Kernel debugging
    • Performance optimisation
    • Operator fusion
  • Experience across both NVIDIA GPU and custom-accelerator ecosystems is preferred
  • Background in compiler engineering, MLIR, or intermediate-representation lowering is advantageous
  • Strong understanding of memory-hierarchy optimisation is preferred
  • Contributions to kernel or accelerator libraries such as cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls are advantageous
  • Strong written communication and ability to provide precise technical feedback
Engagement Details
  • Part-time independent contractor engagement
  • Fully remote within the United States
  • Flexible scheduling based on project requirements
  • Compensation: $60-$80/hour
  • Work focuses on GPU and accelerator kernel development, numerical correctness, performance optimisation, benchmarking, debugging, and technical quality evaluation
  • Projects may be extended, shortened, or concluded based on project needs and performance
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
  • H1-B and STEM OPT support is unavailable for this engagement
About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote | AWS Trainium Kernel Engineer (NKI) — $60–$80/hour
Remote | AWS Trainium Kernel Engineer (NKI) — $60–$80/hour

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 83,000 - 110,000
GPU Kernel Expert Mercor · Remote — United States $70-90/hr →
GPU Kernel Expert Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 100,000 - 180,000
Remote Senior GPU Kernel Evaluator & Optimisation Expert
Remote Senior GPU Kernel Evaluator & Optimisation Expert

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 83,000 - 110,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Obsidian • San Francisco (CA)

On-site
USD 165,000 - 276,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
37303 HD - CUDA Engineering Expert
37303 HD - CUDA Engineering Expert

Cephas Consultancy Services Private Limited • California (MO)

Hybrid
USD 120,000 - 180,000
GPU Kernel Evaluation Expert
GPU Kernel Evaluation Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 150,000 - 350,000