Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)

River AI

Austin (TX)

On-site

USD 200,000 - 420,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Relocation support
Visa sponsorship

Job summary

River AI is building the foundational compute engine for its high-performance silicon in Austin or Palo Alto. You will design kernel generators that emit optimized assembly and push the hardware to its theoretical limits for DL operations like GEMMs and custom activations.

The role requires shipping highly optimized kernels, working with compiler engineers and hardware teams, and optimizing memory and pipelines for maximum throughput.

Qualifications

  • Bachelor's degree in a related field with 5+ years of experience in low-level performance programming.
  • Strong understanding of hardware programming models and shipped kernels.
  • Deep knowledge of computer architecture, vector units, and memory hierarchies.
  • Proficiency in modern C++ for meta-programming and code-generation frameworks.

Responsibilities

  • Design and build C++ code-generation frameworks that emit optimized ISA assembly for custom silicon.
  • Author and optimize core DL primitives (GEMM/MatMul, Attention, Convolutions).
  • Tune instruction scheduling, register allocation, and software pipelines.
  • Design tiling, double-buffering, and data-m movement for on-chip SRAM.
  • Collaborate with RTL/architecture teams to inform ISA and future units.
  • Benchmark against hardware simulators and ensure mathematical correctness.

Skills

C++ performance programming
High-performance computing
Low-level software development
Collaborative cross-team work
Mathematics for DL

Education

Bachelor's degree in Computer Engineering/Computer Science/Electrical Engineering

Tools

CUDA
Triton
CUTLASS
Custom accelerator assembly

Job description

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are looking for exceptional performance and kernel generation engineers to build the foundational compute engine for our high-performance custom silicon. In this role, you will design and implement robust kernel generators that programmatically emit optimized low-level assembly code for our greenfield hardware architecture.

You will bridge the gap between high-level compilation and raw hardware capability, pushing our custom architecture to its absolute theoretical limits for critical deep learning operations (including GEMMs, FlashAttention, and custom activations). You will collaborate closely up and down the stack with compiler engineers, silicon architects, and deep learning researchers to unlock maximum compute efficiency.

What You’ll Do
  • Kernel Generator Development: Design and build C++ code-generation frameworks and meta-programming toolchains that automatically emit optimized custom ISA assembly code.
  • Low-Level Compute Optimization: Author and optimize core deep learning primitives (GEMM/MatMul, Attention mechanisms, Convolutions, and element-wise layers) directly targeted at our custom hardware.
  • Microarchitectural Tuning: Hand-craft and automate instruction scheduling, register allocation, and software pipelining to maximize ALU utilization and hide execution latency on our silicon.
  • Memory Hierarchy Management: Design sophisticated tiling, double-buffering, and data-movement strategies to optimize on-chip SRAM utilization and minimize memory bandwidth bottlenecks.
  • HW/SW Co-Design: Partner with the RTL and architecture teams to evaluate hardware simulations, provide feedback on the ISA, and influence the design of future compute units based on kernel execution profiles.
  • Performance Profiling & Validation: Benchmark generated assembly against hardware simulators and silicon, utilizing hardware performance counters to eliminate performance gaps and ensure mathematical correctness.
Minimum Qualifications:
  • Bachelor’s degree in Computer Engineering, Computer Science, Electrical Engineering, or a related field, and 5+ years of practical industry experience in low-level performance programming.
  • Deep understanding of hardware programming models (e.g., CUDA, Triton, CUTLASS, or custom accelerator assembly) and a proven track record of shipping highly optimized kernels.
  • Advanced knowledge of Computer Architecture, including vector units, execution pipelines, register files, and complex memory hierarchies (caches, SRAM, HBM/DRAM).
  • Proficiency in modern C++ for building robust, scalable meta-programming and code-generation frameworks.
  • Strong mathematical foundation in linear algebra operations and deep learning primitives.
  • A highly collaborative mindset to push boundaries and co-design effectively with hardware and compiler teams.
Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)
  • Deep familiarity with implementing microarchitectural optimizations for Tensor Cores, matrix multiply-accumulate units, or custom vector extensions.
  • Experience utilizing advanced C++ template metaprogramming or code-generation techniques to automate the creation of heavily parameterized kernel variants.
  • Advanced experience with low-level hardware profiling tools, execution tracing, and utilizing performance counters to identify cache misses, pipeline stalls, and ALU bubbles.
Logistics
  • Location: This role is based in Austin, Texas or Palo Alto, California.
  • Compensation: Depending on background, skills, experience, and location, the expected annual salary range for this position is $200,000 - $420,000 USD.
  • Visa Sponsorship: We sponsor visas. We can't guarantee success for every candidate or role, but if you're the right fit, we're committed to working through the visa process.
  • Benefits: River AI offers generous health, dental, and vision benefits, unlimited PTO, and relocation support as needed.

https://job-boards.greenhouse.io/riverai/jobs/4300352009

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)
Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)

River AI Inc. • Austin (TX), Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
Member of Technical Staff, Hardware, Compiler Engineer
Member of Technical Staff, Hardware, Compiler Engineer

River AI Inc. • Austin (TX), Palo Alto (CA)

On-site
USD 200,000 - 420,000
Visa sponsorship
Relocation assistance
Comprehensive benefits
Member of Technical Staff, Hardware, Performance Engineer
Member of Technical Staff, Hardware, Performance Engineer

River AI • San Francisco (CA)

On-site
USD 200,000 - 420,000
Health benefits
Dental and vision benefits
Unlimited PTO
+1
Member of Technical Staff, Hardware, Performance Engineer
Member of Technical Staff, Hardware, Performance Engineer

River AI • Austin (TX)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
Member of Technical Staff, Hardware, Performance Engineer
Member of Technical Staff, Hardware, Performance Engineer

River AI Inc. • Austin (TX), Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health benefits
Dental benefits
Vision benefits
+2
Kernel Engineer, Custom AI Silicon — Compute
Kernel Engineer, Custom AI Silicon — Compute

River AI Inc. • Austin (TX), Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
Member of Technical Staff, Hardware, RTL Design Engineer
Member of Technical Staff, Hardware, RTL Design Engineer

River AI • San Francisco (CA)

On-site
USD 200,000 - 420,000
Health benefits
Dental benefits
Vision benefits
+2
Kernel Engineer for High-Performance AI Compute
Kernel Engineer for High-Performance AI Compute

River AI • Austin (TX)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
+1
Kernel Engineer (Compute / Accelerator)
Kernel Engineer (Compute / Accelerator)

DensityAI • Mountain View (CA)

On-site
USD 260,000 - 320,000
Medical, dental, and vision coverage
401(k) plan
Standard PTO
+1
Principal Software Engineer - Kernels
Principal Software Engineer - Kernels

d-Matrix • Santa Clara (CA)

Hybrid
USD 120,000 - 180,000