Kernel Engineer, Custom AI Silicon — Compute

River AI Inc.

Austin, Palo Alto (TX, CA)

On-site

USD 200,000 - 420,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Relocation support

Job summary

River AI Inc. is hiring a kernel engineer to build the foundational compute engine for its custom silicon. You will design and implement robust kernel generators and emit optimized assembly for our greenfield architecture.

You will collaborate with compiler engineers, silicon architects, and DL researchers to push compute efficiency, optimize GEMMs and attention ops, and ensure correctness on hardware simulations and silicon.

Qualifications

  • Bachelor's degree in Computer Engineering, CS, or Electrical Engineering with 5+ years of industry experience in low-level performance programming.
  • Deep understanding of hardware programming models (CUDA, Triton, CUTLASS) and a proven track record of shipping highly optimized kernels.
  • Advanced knowledge of Computer Architecture, including vector units, execution pipelines, register files, and memory hierarchies (caches, SRAM, HBMs).
  • Proficiency in modern C++ for building robust, scalable meta-programming and code-generation frameworks.
  • Strong mathematical foundation in linear algebra and deep learning primitives.
  • A highly collaborative mindset to push boundaries and co-design effectively with hardware and compiler teams.

Responsibilities

  • Kernel Generator Development: Design and build C++ code-generation frameworks and meta-programming toolchains that automatically emit optimized custom ISA assembly code.
  • Low-Level Compute Optimization: Author and optimize core deep learning primitives (GEMM/MatMul, Attention, Convolutions, and element-wise layers) for custom hardware.
  • Microarchitectural Tuning: Hand-craft and automate instruction scheduling, register allocation, and software pipelining to maximize ALU utilization and hide latency.
  • Memory Hierarchy Management: Design tiling, double-buffering, and data-movement strategies to maximize on-chip SRAM usage and minimize memory bottlenecks.
  • HW/SW Co-Design: Partner with RTL/architecture teams to evaluate hardware simulations, provide ISA feedback, and influence future compute units.
  • Performance Profiling & Validation: Benchmark generated assembly against hardware simulators and silicon; use performance counters to close gaps and ensure correctness.

Skills

Low-level performance programming
C++ metaprogramming
Hardware programming models
GEMM/MatMul
Tensor Core optimization
Collaboration with hardware/compiler

Education

Bachelor's degree in Computer Engineering/CS/EE

Tools

CUDA
Triton
CUTLASS
Custom ISA assembly

Job description

River AI Inc. is hiring a kernel engineer to build the foundational compute engine for its custom silicon. You will design and implement robust kernel generators and emit optimized assembly for our greenfield architecture.

You will collaborate with compiler engineers, silicon architects, and DL researchers to push compute efficiency, optimize GEMMs and attention ops, and ensure correctness on hardware simulations and silicon.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kernel Engineer for High-Performance AI Compute
Kernel Engineer for High-Performance AI Compute

River AI • Austin (TX)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
+1
Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)
Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)

River AI • Austin (TX)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
+1
Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)
Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)

River AI Inc. • Austin (TX), Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
AI Compiler Engineer, Custom Silicon
AI Compiler Engineer, Custom Silicon

River AI Inc. • Austin (TX), Palo Alto (CA)

On-site
USD 200,000 - 420,000
Visa sponsorship
Relocation assistance
Comprehensive benefits
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip
Kernel Engineer - High-Performance ML/HPC on Custom AI Chip

Cerebras • United States

On-site
USD 100,000 - 130,000
Opportunity to publish open-source AI research
Work with one of the fastest AI supercomputers
Non-corporate work culture
New Grad Kernel Engineer for AI HPC Systems
New Grad Kernel Engineer for AI HPC Systems

Cerebras • United States

On-site
USD 120,000 - 180,000
Kernel Engineer for AI Accelerator - Profiling
Kernel Engineer for AI Accelerator - Profiling

DensityAI • Mountain View (CA)

On-site
USD 260,000 - 320,000
Medical, dental, and vision coverage
401(k) plan
Standard PTO
+1
Remote Kernel Engineer for AI Hardware & ML Kernels
Remote Kernel Engineer for AI Hardware & ML Kernels

United States Digital Space LLC • United States

Remote
USD 140,000 - 200,000
AI Hardware Kernel Performance Engineer
AI Hardware Kernel Performance Engineer

Cerebras • Sterling (VA)

On-site
USD 100,000 - 150,000
Senior Hardware Performance Engineer for AI Accelerators
Senior Hardware Performance Engineer for AI Accelerators

River AI • San Francisco (CA)

On-site
USD 200,000 - 420,000
Health benefits
Dental and vision benefits
Unlimited PTO
+1