AWS Trainium / NKI Kernel Expert

Anyone AI

United States

Remote

USD 83,000 - 138,000

Part time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

anyone-ai is seeking a remote, part-time consultant to evaluate and refine AWS Trainium and Neuron Kernel Interface tasks for AI workloads. The role requires deep kernel experience and low-level optimization.

You will review NKI kernel correctness, CUDA-to-NKI migrations, memory hierarchies, and DMA scheduling, delivering structured technical feedback and opportunities for optimization.

Qualifications

  • Two or more years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface.
  • Experience with AWS Trainium and/or Inferentia2 hardware.
  • Strong grasp of tile-based computation, SBUF/PSUM/HBM memory hierarchy, partition dimension constraints, and DMA orchestration.
  • Ability to evaluate CUDA-to-NKI migrations critically.
  • Background profiling and optimizing workloads on Trainium.
  • Understanding of numerical differences across GPU and Trainium backends.

Responsibilities

  • Review NKI kernel correctness and Trainium-specific development patterns.
  • Evaluate CUDA-to-NKI kernel migrations for idiomatic Trainium implementation.
  • Assess performance optimization and benchmarking on Trainium hardware.
  • Review memory management across the SBUF, PSUM, and HBM hierarchy.
  • Analyze tile-based computation and DMA scheduling decisions.
  • Evaluate cross-platform numerical correctness between CUDA/Triton and NKI.
  • Identify Trainium-specific bottlenecks and surface optimization opportunities.
  • Provide structured written technical feedback and quality assessments.

Skills

NKI kernel development
Trainium hardware experience
Memory hierarchy optimization
CUDA-to-NKI migrations evaluation
Performance profiling

Tools

AWS Neuron SDK
Neuron Compiler

Job description

Role overview

This remote, part-time consulting role focuses on evaluating and refining AWS Trainium and Neuron Kernel Interface tasks for AI workloads. The work requires deep familiarity with Trainium architecture and an understanding of how it diverges from conventional GPU programming paradigms.

Responsibilities
  • Review NKI kernel correctness and Trainium-specific development patterns
  • Evaluate CUDA-to-NKI kernel migrations for idiomatic Trainium implementation
  • Assess performance optimization and benchmarking on Trainium hardware
  • Review memory management across the SBUF, PSUM, and HBM hierarchy
  • Analyze tile-based computation and DMA scheduling decisions
  • Evaluate cross-platform numerical correctness between CUDA/Triton and NKI
  • Identify Trainium-specific bottlenecks and surface optimization opportunities
  • Provide structured written technical feedback and quality assessments
Requirements
  • Two or more years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface
  • Experience working with AWS Trainium and/or Inferentia2 hardware
  • Strong grasp of tile-based computation, SBUF/PSUM/HBM memory hierarchy, partition dimension constraints, and DMA orchestration
  • Ability to evaluate CUDA-to-NKI migrations critically
  • Background profiling and optimizing workloads on Trainium
  • Understanding of numerical differences across GPU and Trainium backends
Nice to have
  • Experience with the AWS Neuron SDK or Neuron Compiler
  • CUDA or Triton kernel development background
  • Familiarity with NeuronCore-v2 architecture
  • Working with FP32, BF16, FP8, or INT8 workloads
  • Benchmarking on Trn1 or Trn2 instances
  • Knowledge of nki.language, @nki.jit, or XLA custom calls
  • Background in technical evaluation, RLHF, or rubric-based assessment
Benefits and work setup

Remote, part-time, project-based consulting engagement focused on AWS Trainium and NKI kernel engineering and evaluation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

NKI Kernel Optimization Specialist
NKI Kernel Optimization Specialist

Obsidian • San Francisco (CA)

On-site
USD 150,000 - 210,000
NKI Kernel Optimization Specialist
NKI Kernel Optimization Specialist

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
Trainium NKI Kernel Expert
Trainium NKI Kernel Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Contractor role
Part-time opportunity
US-based role
+3
Trainium Kernel Expert - AI Trainer
Trainium Kernel Expert - AI Trainer

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
Trainium Kernel Expert - AI Trainer
Trainium Kernel Expert - AI Trainer

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 180,000
Trainium (NKI) Kernel Expert
Trainium (NKI) Kernel Expert

DigiNo • Northern (KY)

Hybrid
USD 150,000 - 210,000
Trainium NKI Kernel Expert — Remote Part-Time
Trainium NKI Kernel Expert — Remote Part-Time

Anyone AI • United States

On-site
USD 83,000 - 138,000
Trainium Kernel Engineer - NKI AI Trainer
Trainium Kernel Engineer - NKI AI Trainer

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
NKI Kernel Expert - Trainium Performance & Migration
NKI Kernel Expert - Trainium Performance & Migration

DigiNo • Northern (KY)

Hybrid
USD 150,000 - 210,000
Senior Kernel Engineer: 26-02794
Senior Kernel Engineer: 26-02794

Akraya, Inc. • Bellevue (WA)

On-site
USD 117,000 - 124,000