NPU Kernel/Operator Engineer

Black Sesame Technologies Inc

San Jose (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Black Sesame Technologies Inc is seeking a Senior NPU Kernel/Operator Engineer in San Jose, California. This role focuses on designing and optimizing high-performance kernels for a custom AI accelerator, involving deep learning operations and performance optimization.

The ideal candidate will have over 5 years of experience in accelerator programming and a strong foundation in memory hierarchy and performance-critical kernel optimization. Collaboration with hardware and model teams is essential.

Qualifications

  • 5+ years of experience in performance optimization and accelerator programming.
  • Deep understanding of memory hierarchy, tiling, parallelism, and bandwidth analysis.
  • Experience optimizing performance-critical kernels or numerical computation.

Responsibilities

  • Design and optimize high-performance NPU kernels for neural network workloads.
  • Own critical operators such as attention-style kernels and normalization.
  • Optimize data movement across various memory architectures.

Skills

Performance optimization
Accelerator programming
GPU development
NPU programming
HPC systems

Education

BS/MS/PhD in CS, EE, Computer Engineering, or related field

Tools

CUDA
Triton
OpenCL

Job description

We are looking for a Senior NPU Kernel/Operator Engineer to lead the design and optimization of high-performance kernels for a custom AI accelerator / NPU. This role focuses on general-purpose deep learning operators, fused kernels, and hardware-aware performance optimization across CNNs, transformers, and other neural network workloads.

The ideal candidate has strong experience in performance engineering on GPU, NPU, DSP, CPU SIMD, compiler backend, embedded accelerator, or HPC systems.

Responsibilities
  • Design and optimize high-performance NPU kernels for a broad range of neural network workloads.
  • Own critical operators such as attention-style kernels, normalization, reduction, layout conversion, gather/scatter, quant/dequant, and fused operators.
  • Develop tiling, blocking, vectorization, and memory scheduling strategies.
  • Optimize data movement across matrix engine, vector engine, SRAM, DMA, NoC, cache, and DRAM.
  • Analyze bottlenecks in compute utilization, memory bandwidth, synchronization, DMA overlap, bank conflicts, and instruction overhead.
  • Build first-principles performance models for key operators.
  • Drive kernels toward hardware roofline limits.
  • Collaborate with hardware, compiler, runtime, and model teams on ISA features, tensor layouts, memory access patterns, and operator APIs.
  • Debug complex correctness, precision, and performance issues on simulator or silicon.
  • Mentor junior engineers and establish kernel optimization best practices.
Requirements
  • BS/MS/PhD in CS, EE, Computer Engineering, or related field.
  • 5+ years of experience in performance optimization, accelerator programming, GPU/NPU/DSP development, compiler backend, embedded systems, or HPC.
  • Deep understanding of memory hierarchy, tiling, parallelism, vectorization, synchronization, and bandwidth analysis.
  • Experience optimizing performance-critical kernels or numerical computation.
  • Ability to reason from algorithm requirements to hardware execution and performance bottlenecks.
Preferred
  • Experience with CUDA, Triton, CUTLASS, OpenCL, TVM, MLIR, Halide, SIMD intrinsics, DSP SDKs, or custom accelerator SDKs.
  • Experience optimizing operators such as convolution, GEMM, attention, softmax, normalization, reduction, image processing, or fused compute/memory kernels.
  • Familiarity with custom AI accelerator architecture, matrix engines, vector engines, systolic arrays, DMA, SRAM, NoC, or DRAM systems.
  • Experience with mixed precision and quantization: FP32, FP16, BF16, FP8, INT8, INT4.
  • Experience with simulator/emulator/FPGA/silicon bring-up is a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior NPU Kernel & Operator Performance Architect
Senior NPU Kernel & Operator Performance Architect

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 120,000 - 160,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Lead Kernel Engineer/Architect (m/f/d)
Lead Kernel Engineer/Architect (m/f/d)

EPAM Systems • Germany (OH)

Hybrid
USD 104,000 - 152,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs
Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000