Kernel Engineer (Compute / Accelerator)

DensityAI

Mountain View (WY)

On-site

USD 260,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity grant
Medical/dental/vision
401(k)
PTO

Job summary

DensityAI is seeking a software engineer to write, evaluate, and profile specialized compute kernels that run on a custom AI accelerator. You’ll work with architecture and compiler teams to define the kernel programming model and drive performance profiling that informs silicon design.

You’ll write and optimize tensor operations and data movement patterns, develop profiling tools, and contribute to MLIR/LLVM related components.

Qualifications

  • Production-grade C/C++ systems code experience.
  • Deep GPU kernel experience with CUDA or equivalent.
  • Strong computer architecture knowledge.
  • Proven performance profiling and optimization skills.
  • Experience mapping tensor operations to hardware.

Responsibilities

  • Write and optimize compute kernels for an AI accelerator.
  • Develop and maintain kernel profiling infrastructure.
  • Define shuffle patterns for ML kernel primitives.
  • Drive kernel DSL design and memory management strategies.
  • Enable end-to-end kernel execution on the architectural simulator.
  • Collaborate on the MLIR/LLVM compiler infrastructure and kernel validation.
  • Create onboarding documentation and kernel writing guides.

Skills

C/C++
CUDA
Computer architecture
Profiling
Tensor operations
Python

Tools

CUTLASS
MLIR
LLVM

Job description

About the role

You will write, evaluate, and profile specialized compute kernels that run on a custom AI accelerator. This is the critical interface between high-level ML workloads and silicon — your code directly determines how effectively the hardware performs. You’ll work closely with the architecture and compiler teams to define the kernel programming model, implement core tensor operations, and drive the performance profiling workflow that validates silicon design decisions.

What you’ll do
  • Write and optimize compute kernels for a custom AI accelerator — tensor operations, data movement patterns, memory hierarchy exploitation
  • Develop and maintain profiling infrastructure to measure kernel performance against architectural targets
  • Define and document shuffle patterns for ML kernel primitives across CPU‑like control, tensor cores, and CUTLASS‑style operations
  • Drive kernel DSL design decisions — thread spawn mechanisms, register passing conventions, and memory management strategies
  • Enable end‑to‑end kernel execution on the architectural simulator
  • Collaborate with the compiler team on the MLIR dialect — your kernels are the primary validation target
  • Create onboarding documentation and kernel writing guides for the broader team
What we’re looking for
  • C/C++ — production‑grade systems code, not scripted glue. You’ll write performance‑critical kernels
  • CUDA or equivalent accelerator programming — deep experience writing GPU kernels, understanding warp/wavefront execution, memory coalescing, shared memory optimization. The mental model transfers directly
  • Computer architecture — you need to reason about pipelines, memory hierarchies, data movement costs, and how software maps to hardware
  • Performance profiling and optimization — you live in profilers. Identifying bottlenecks, measuring throughput, and iterating until kernels meet targets is the core loop
  • Tensor operations — practical understanding of GEMM, convolution, attention, reduction, and scatter/gather as they map to hardware
  • Python — for scripting, DSL integration, and profiling automation
  • (Optional) RISC‑V, x86, or ARM64 ISA experience
  • (Optional) MLIR or LLVM compiler infrastructure
  • (Optional) HPC or scientific computing background (large‑scale parallel compute intuition)
  • (Optional) FPGA or Verilog/SystemVerilog (ability to read RTL and reason about the hardware you’re targeting)
  • (Optional) Familiarity with CUTLASS, Triton, or similar kernel libraries
Compensation

Final offers depend on level, location, and skills relevant to the role. Additional compensation: equity grant per company guidelines; medical / dental / vision; 401(k); standard PTO.

Visa Sponsorship

DensityAI sponsors qualified candidates for H‑1B, O‑1, TN, E‑3, and other employment‑based visas, and we welcome applicants on F‑1 OPT and STEM‑OPT. Work authorization is required at start; we provide immigration support to secure or transfer status.

Equal Opportunity

DensityAI is an Equal Opportunity Employer. We do not discriminate on the basis of race, color, religious creed, national origin, ancestry, physical or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, age (40+), sexual orientation, military or veteran status, pregnancy, or any other status protected by law. We comply with the California CROWN Act and provide reasonable accommodations on request.

Full compensation packages are based on candidate experience and relevant certifications.

California pay range

$260,000—$320,000 USD

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kernel Engineer (Compute / Accelerator)
Kernel Engineer (Compute / Accelerator)

DensityAI • Mountain View (CA)

On-site
USD 260,000 - 320,000
Medical, dental, and vision coverage
401(k) plan
Standard PTO
+1
Compiler Engineer — LLVM Backend
Compiler Engineer — LLVM Backend

DensityAI • Mountain View (CA)

On-site
USD 180,000 - 320,000
Equity grant
Medical/Dental/Vision
401(k)
+1
Kernel Engineer for AI Accelerator - Profiling
Kernel Engineer for AI Accelerator - Profiling

DensityAI • Mountain View (CA)

On-site
USD 260,000 - 320,000
Medical, dental, and vision coverage
401(k) plan
Standard PTO
+1
Performance Verification Engineer
Performance Verification Engineer

DensityAI • Mountain View (WY)

On-site
USD 200,000 - 350,000
Equity grant
Medical / dental / vision
401(k)
+1
Compiler Engineer — MLIR
Compiler Engineer — MLIR

DensityAI • Mountain View (CA)

On-site
USD 200,000 - 360,000
Equity grant
Medical benefits
401(k)
+1
Kernel Engineer, AI Accelerator Performance & Optimization
Kernel Engineer, AI Accelerator Performance & Optimization

DensityAI • Mountain View (WY)

On-site
USD 260,000 - 320,000
Equity grant
Medical/dental/vision
401(k)
+1
Electrical Engineer, Hardware Systems
Electrical Engineer, Hardware Systems

DensityAI • Mountain View (WY)

On-site
USD 200,000 - 320,000
Equity grant
401(k)
Medical/dental/vision
+1
Performance Verification Engineer
Performance Verification Engineer

DensityAI • Mountain View (CA)

On-site
USD 200,000 - 350,000
Equity grant
Medical/dental/vision insurance
401(k) plan
+1
Principal Software Engineer – Kernels
Principal Software Engineer – Kernels

MixMode • Santa Clara (CA)

Hybrid
USD 150,000 - 200,000
Principal Software Engineer - Kernels
Principal Software Engineer - Kernels

d-Matrix • Santa Clara (CA)

Hybrid
USD 120,000 - 180,000