Senior Kernel Engineer: 26-02794

Akraya, Inc.

Bellevue (WA)

On-site

USD 117,000 - 124,000

Full time

9 days ago
Application generator

Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Akraya, Inc. is seeking an experienced kernel optimization engineer to develop and optimize low-level compute kernels for AI accelerators in Bellevue, WA.

You will ensure training environments reflect real-world kernel engineering challenges by writing, debugging, and tuning operators, while using profiling tools and accelerator interfaces to validate task realism. The role requires hands-on work with CUDA/Triton or NKI and strong coding skills, with a background in high-performance compute and

Qualifications

  • 7+ years of relevant experience required.
  • Experience with AWS Trainium/Inferentia and NKI.
  • Hands-on proficiency with CUDA, Triton, or NKI.
  • Familiarity with Neuron SDK, Neuron Profiler, and framework integration (PyTorch, JAX).
  • Experience porting kernels between accelerator backends (e.g., CUDA/Triton to NKI).

Responsibilities

  • Develop and optimize custom compute kernels for AI accelerators.
  • Validate training environment tasks for realistic kernel development scenarios.
  • Debug and profile kernel code to identify and resolve performance issues.
  • Analyze kernel performance against hardware constraints and memory hierarchy.
  • Consult with training environment teams on task realism and quality.

Skills

CUDA programming
Triton optimization
Profiling tools
Debugging kernels
Performance tuning

Education

Bachelor's degree in CS/CE/EE or equivalent

Tools

AWS Trainium/Inferentia
Neuron Kernel Interface (NKI)
Neuron SDK
Neuron Profiler
PyTorch
JAX

Job description

Primary Skills:
  • CUDA programming (advanced)
  • Triton optimization (advanced)
  • Profiling tools (advanced)
  • Debugging kernels (advanced)
  • Performance tuning (advanced)

Contract Type: W2

Duration: 3+ months with possible extension or conversion

Location: Bellevue, WA (#LI-Onsite)

Pay Range: $85.00 - $90.00 Per hour on W2 #LP

Job Summary

This role involves the low-level development and optimization of compute kernels for AI accelerator hardware. You will ensure training environments accurately mirror real-world kernel engineering challenges by writing, debugging, and optimizing operators. A key aspect of this position is using profiling tools and accelerator programming interfaces to validate task realism and define solution quality.

Key Responsibilities
  • Develop and optimize custom compute kernels for AI accelerators.
  • Validate training environment tasks for realistic kernel development scenarios.
  • Debug and profile kernel code to identify and resolve performance issues.
  • Analyze kernel performance against hardware constraints and memory hierarchy.
  • Consult with training environment development teams on task realism and quality.
Must-Have Skills
  • Computer Science, Computer Engineering, Electrical Engineering, or equivalent technical degree and 7+ years of relevant experience is required.
  • Direct experience with AWS Trainium/Inferentia and the Neuron Kernel Interface (NKI) and Hands-on proficiency with a low-level kernel programming language such as CUDA, Triton, or the AWS Neuron Kernel Interface (NKI).
  • Familiarity with the Neuron SDK, Neuron Profiler, and framework integration (PyTorch, JAX).
  • Experience porting kernels between accelerator backends (e.g., CUDA/Triton to NKI) and understanding of deep learning operator implementations (attention, GEMM, normalization).
Industry Experience

Experience in developing and optimizing low-level code for AI accelerators is essential, with a strong preference for experience with Triton over CUDA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance
GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

anyone-ai • United States

Remote
USD 138,000 - 248,000
AWS Trainium / NKI Kernel Expert
AWS Trainium / NKI Kernel Expert

Anyone AI • United States

Remote
USD 83,000 - 138,000
Senior AI Kernel Engineer – CUDA/Triton Optimizer
Senior AI Kernel Engineer – CUDA/Triton Optimizer

Akraya, Inc. • Bellevue (WA)

On-site
USD 117,000 - 124,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Obsidian • Chicago (IL)

On-site
USD 83,000 - 165,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
NKI Kernel Optimization Specialist
NKI Kernel Optimization Specialist

Obsidian • San Francisco (CA)

On-site
USD 150,000 - 210,000
NKI Kernel Optimization Specialist
NKI Kernel Optimization Specialist

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
NPU Kernel/Operator Engineer
NPU Kernel/Operator Engineer

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs
Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000