Remote | AWS Trainium Kernel Engineer (NKI) — $60–$80/hour

24-Mag Llc

New York, Northern (NY, KY)

Hybrid

USD 83,000 - 110,000

Part time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

24-MAG LLC is offering a specialised part-time consulting opportunity for experienced kernel engineers with hands-on Neuron Kernel Interface (NKI) expertise. You will focus on evaluating NKI kernel development tasks for technical correctness, hardware appropriateness, numerical fidelity, and performance quality in a fully remote capacity across the United States.

The role covers CUDA-to-NKI migrations, tile-based computation, memory hierarchy management (SBUF, PSUM, HBM), DMA orchestration, and

Qualifications

  • 2+ years hands-on kernel experience with the Neuron Kernel Interface (NKI).
  • Experience targeting AWS Trainium or Inferentia2 hardware.
  • Experience with CUDA-to-NKI migrations.
  • Strong understanding of tile-based computation and memory hierarchies.
  • Ability to provide rubric-based technical feedback and precise recommendations.

Responsibilities

  • Evaluate NKI kernels for correctness and hardware suitability.
  • Review CUDA kernels migrated to NKI for semantic preservation.
  • Assess tile-based computation and memory movement strategies.
  • Evaluate Trainium-specific performance and profiling evidence.
  • Provide clear, rubric-based written feedback and recommendations.

Skills

NKI Experience
Trainium Hardware
Kernel Development
CUDA Migration
Performance Optimization
NKI Libraries

Tools

AWS Neuron SDK
Neuron Compiler
NKI Tooling

Job description

We are sharing a specialised part-time consulting opportunity for experienced kernel engineers with hands‑on expertise in the Neuron Kernel Interface (NKI), AWS Trainium/Inferentia2 hardware, low‑level performance optimisation, and CUDA‑to‑NKI migration.

This role focuses on evaluating NKI kernel‑development tasks for technical correctness, hardware appropriateness, numerical fidelity, and performance quality. Selected experts will review Trainium‑specific implementations, migration decisions, memory‑management strategies, profiling results, and cross‑platform numerical behaviour while providing clear, rubric‑based technical feedback.

Key Responsibilities
NKI Kernel Development Review
  • Evaluate kernels developed using the Neuron Kernel Interface (NKI)
  • Assess whether implementations appropriately target AWS Trainium and Inferentia2 hardware
  • Review low‑level computation patterns for technical correctness
  • Identify inefficient, incorrect, or hardware‑inappropriate implementation choices
  • Apply practical judgement grounded in hands‑on NKI development experience
CUDA‑to‑NKI Migration
  • Review migrations of existing CUDA kernels to NKI
  • Assess whether computational semantics are preserved across platforms
  • Identify translation errors, unsupported assumptions, or inefficient migration strategies
  • Evaluate whether NKI implementations appropriately account for Trainium architecture
  • Distinguish faithful migrations from implementations that merely reproduce surface‑level CUDA structure
Tile‑Based Computation
  • Assess tile decomposition and computation strategies
  • Review partitioning decisions against NKI execution constraints
  • Evaluate whether kernels make effective use of available compute resources
  • Identify inefficient tiling or data‑movement patterns
  • Assess whether implementation choices align with NKI programming requirements
Memory Hierarchy Management
  • Review use of SBUF, PSUM, and HBM
  • Assess data placement and movement across Trainium memory hierarchies
  • Evaluate memory‑bandwidth utilisation and locality
  • Identify unnecessary transfers or memory bottlenecks
  • Review implementation decisions affecting on‑chip memory efficiency
DMA & Data Movement
  • Evaluate DMA orchestration within NKI kernels
  • Review sequencing of computation and data‑transfer operations
  • Identify stalls, inefficient transfer patterns, or synchronisation issues
  • Assess whether data movement appropriately overlaps with computation
  • Evaluate implementation choices affecting pipeline utilisation
Trainium Performance Optimisation
  • Review Trainium‑specific optimisation strategies
  • Assess NeuronCore pipeline utilisation, tensor‑engine throughput, and memory‑bandwidth behaviour
  • Identify performance bottlenecks within kernel implementations
  • Evaluate whether optimisation decisions are supported by profiling evidence
  • Review trade‑offs affecting throughput, latency, and resource utilisation
Numerical Correctness
  • Evaluate numerical consistency between GPU and Trainium implementations
  • Review differences caused by accumulation order, rounding behaviour, and mixed‑precision semantics
  • Assess appropriate tolerances for cross‑platform comparisons
  • Identify numerical discrepancies that indicate implementation defects
  • Distinguish expected hardware‑level variation from substantive correctness problems
Precision & Data Types
  • Review kernels using supported formats such as FP32, BF16, FP8, and INT8
  • Assess precision choices against computational requirements
  • Evaluate mixed‑precision behaviour and numerical stability
  • Identify inappropriate casting or accumulation strategies
  • Review whether performance gains are achieved without compromising required correctness
AWS Neuron Ecosystem
  • Evaluate implementations using the AWS Neuron SDK
  • Review interactions between kernel code, compilation, and Trainium execution
  • Assess compiler‑related behaviours where relevant
  • Apply familiarity with NKI kernel libraries and Neuron tooling
  • Identify implementation issues arising from platform‑specific constraints
Benchmarking & Validation
  • Review benchmark results for Trainium workloads
  • Assess performance comparisons and experimental methodology
  • Evaluate workloads running on Trn1 or Trn2 instances where applicable
  • Determine whether claimed performance improvements are supported by evidence
  • Identify benchmarking methodologies that could produce misleading conclusions
Rubric‑Based Technical Evaluation
  • Assess assigned kernel‑development tasks against structured technical criteria
  • Provide clear written explanations supporting evaluation decisions
  • Reference specific implementation, profiling, or numerical evidence
  • Apply evaluation standards consistently across assignments
  • Distinguish valid optimisation alternatives from technically flawed approaches
Ideal Profile
  • 2+ years of hands‑on experience developing or optimising kernels using the Neuron Kernel Interface (NKI)
  • Professional experience targeting AWS Trainium or Inferentia2 hardware
  • Strong understanding of tile‑based computation
  • Deep familiarity with SBUF, PSUM, and HBM memory management
  • Strong knowledge of partition‑dimension constraints and DMA orchestration
  • Demonstrated experience evaluating or performing CUDA‑to‑NKI migrations
  • Familiarity with Trainium‑specific performance profiling
  • Experience assessing NeuronCore pipeline utilisation, tensor‑engine throughput, and memory‑bandwidth bottlenecks
  • Strong understanding of cross‑platform numerical correctness and mixed‑precision behaviour
  • Direct experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel libraries is preferred
  • Prior CUDA or Triton kernel development experience is advantageous
  • Familiarity with NeuronCore‑v2 architecture and supported numerical formats is preferred
  • Experience benchmarking ML workloads on Trn1 or Trn2 instances is advantageous
  • Strong written communication and ability to provide precise technical feedback
Engagement Details
  • Part‑time independent contractor engagement
  • Fully remote within the United States
  • Flexible scheduling based on project requirements
  • Compensation: $60–$80/hour
  • Work focuses on NKI kernel development, Trainium performance optimisation, CUDA migration, numerical correctness, and technical quality evaluation
  • Projects may be extended, shortened, or concluded based on project needs and performance
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
  • H1‑B and STEM OPT support is unavailable for this engagement
About the Platform

This opportunity is available through 24‑MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project‑based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24‑MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Trainium (NKI) Kernel Expert Mercor · Remote — United States $70-90/hr →
Trainium (NKI) Kernel Expert Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 120,000 - 180,000
Trainium NKI Kernel Expert
Trainium NKI Kernel Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Contractor role
Part-time opportunity
US-based role
+3
NKI Kernel Optimization Specialist
NKI Kernel Optimization Specialist

Obsidian • San Francisco (CA)

On-site
USD 150,000 - 210,000
Remote | GPU Kernel Engineer — $60–$80/hour
Remote | GPU Kernel Engineer — $60–$80/hour

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 83,000 - 110,000
NKI Kernel Optimization Specialist
NKI Kernel Optimization Specialist

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
Part-Time NKI Kernel Engineer: Trainium & CUDA Migration
Part-Time NKI Kernel Engineer: Trainium & CUDA Migration

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 83,000 - 110,000
Trainium Kernel Expert - AI Trainer
Trainium Kernel Expert - AI Trainer

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
Trainium Kernel Expert - AI Trainer
Trainium Kernel Expert - AI Trainer

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 180,000
Trainium Kernel Engineer - NKI AI Trainer
Trainium Kernel Engineer - NKI AI Trainer

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
Remote Trainium NKI Kernel Expert: CUDA Migrations
Remote Trainium NKI Kernel Expert: CUDA Migrations

Appsierra Group • United States

On-site
USD 146,000 - 187,000