Kernel Engineer, AI Accelerator Performance & Optimization

DensityAI

Mountain View (WY)

On-site

USD 260,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity grant
Medical/dental/vision
401(k)
PTO

Job summary

DensityAI is seeking a software engineer to write, evaluate, and profile specialized compute kernels that run on a custom AI accelerator. You’ll work with architecture and compiler teams to define the kernel programming model and drive performance profiling that informs silicon design.

You’ll write and optimize tensor operations and data movement patterns, develop profiling tools, and contribute to MLIR/LLVM related components.

Qualifications

  • Production-grade C/C++ systems code experience.
  • Deep GPU kernel experience with CUDA or equivalent.
  • Strong computer architecture knowledge.
  • Proven performance profiling and optimization skills.
  • Experience mapping tensor operations to hardware.

Responsibilities

  • Write and optimize compute kernels for an AI accelerator.
  • Develop and maintain kernel profiling infrastructure.
  • Define shuffle patterns for ML kernel primitives.
  • Drive kernel DSL design and memory management strategies.
  • Enable end-to-end kernel execution on the architectural simulator.
  • Collaborate on the MLIR/LLVM compiler infrastructure and kernel validation.
  • Create onboarding documentation and kernel writing guides.

Skills

C/C++
CUDA
Computer architecture
Profiling
Tensor operations
Python

Tools

CUTLASS
MLIR
LLVM

Job description

DensityAI is seeking a software engineer to write, evaluate, and profile specialized compute kernels that run on a custom AI accelerator. You’ll work with architecture and compiler teams to define the kernel programming model and drive performance profiling that informs silicon design.

You’ll write and optimize tensor operations and data movement patterns, develop profiling tools, and contribute to MLIR/LLVM related components.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kernel Engineer for AI Accelerator - Profiling
Kernel Engineer for AI Accelerator - Profiling

DensityAI • Mountain View (CA)

On-site
USD 260,000 - 320,000
Medical, dental, and vision coverage
401(k) plan
Standard PTO
+1
Kernel Engineer (Compute / Accelerator)
Kernel Engineer (Compute / Accelerator)

DensityAI • Mountain View (CA)

On-site
USD 260,000 - 320,000
Medical, dental, and vision coverage
401(k) plan
Standard PTO
+1
Kernel Engineer (Compute / Accelerator)
Kernel Engineer (Compute / Accelerator)

DensityAI • Mountain View (WY)

On-site
USD 260,000 - 320,000
Equity grant
Medical/dental/vision
401(k)
+1
Kernel Engineer for High-Performance AI Compute
Kernel Engineer for High-Performance AI Compute

River AI • Austin (TX)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
+1
Kernel Engineer, Custom AI Silicon — Compute
Kernel Engineer, Custom AI Silicon — Compute

River AI Inc. • Austin (TX), Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip

Cerebras Systems • United States

On-site
USD 110,000 - 140,000
Non-corporate work culture
Equal opportunity employer
Continuous learning and growth opportunities
AI Systems Engineer: High-Performance ML on Accelerators
AI Systems Engineer: High-Performance ML on Accelerators

AMD • San Jose (CA)

On-site
USD 150,000 - 190,000
Senior AI Kernel & Performance Engineer | Equity
Senior AI Kernel & Performance Engineer | Equity

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Kernel Driver Architect for AI Accelerators
Kernel Driver Architect for AI Accelerators

The Consensus • San Jose (CA)

On-site
USD 180,000 - 280,000
Medical, dental, and vision benefits
Relocation to San Jose
Housing subsidy
+3
AI Systems Engineer: High-Performance ML on Accelerators
AI Systems Engineer: High-Performance ML on Accelerators

Advanced Micro Devices • San Jose (CA)

On-site
USD 140,000 - 190,000