Compiler Engineer, Hardware

River AI

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Relocation support

Job summary

River AI in Palo Alto or Austin is seeking an exceptional AI compiler engineer to bridge PyTorch graphs to our custom ISA. You will own kernel algorithms, IRs, and even ISA tweaks, collaborating with researchers and silicon architects.

You will design MLIR-based lowering passes, implement backends for our hardware, and optimize performance across the stack. The role requires a strong CS background and teamwork.

Qualifications

  • Bachelor's degree in Electrical Engineering or Computer Engineering with 5+ years of relevant industry experience.
  • Deep hands-on experience with MLIR or XLA for deep learning workloads.
  • Expert-level understanding of PyTorch internals and integration with external backends.
  • Proficiency in modern C/C++ for high-performance compiler infrastructure.
  • Strong knowledge of computer architecture and various chip programming models.

Responsibilities

  • Lower PyTorch models into custom hardware using MLIR dialects and LLVM-based tooling.
  • Develop and maintain the backend toolchain for our silicon including scheduling and register allocation.
  • Design tiling and fusion strategies to maximize bandwidth and minimize memory movement.
  • Integrate high-performance kernels (Triton/CUDA-like) into the compiler flow.
  • Collaborate with RTL, architecture, and performance teams to optimize ISA and compiler outcomes.

Skills

PyTorch internals
Collaborative mindset
Advanced computer architecture

Education

Bachelor’s degree in Electrical Engineering or Computer Engineering

Tools

MLIR
XLA
C/C++
LLVM-based toolchain

Job description

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are looking for exceptional AI compiler engineers to build the software bridge between the newest AI models and our high-performance custom silicon. You will create and build the compiler stack from PyTorch graphs all the way to optimized custom ISA assembly code. You will take ownership of kernel algorithms, intermediate representations, and even modify the ISA as necessary to achieve a flexible and high performance compiler stack. You will be collaborating both up and down the stack with AI researchers and modelers, as well as with performance engineers and silicon architects.

What You’ll Do
  • Graph Lowering & Optimization: Design and implement compiler passes to lower PyTorch models into custom hardware, leveraging MLIR dialects and LLVM frameworks.
  • Custom Backend Development: Develop and maintain the backend toolchain for our custom silicon, including instruction scheduling, register allocation, and hardware-specific code generation.
  • Memory & Loop Transformations: Design sophisticated tiling and fusion strategies to maximize bandwidth utilization and minimize on-chip memory movement.
  • Kernel Integration: Collaborate with software and hardware teams to integrate high-performance kernels (Triton/CUDA-like) into the automated compiler flow.
  • Performance Profiling: Identify "compilation gaps" where the compiler fails to achieve peak hardware performance, and collaborate with the performance team for targeted optimizations to close those gaps.
  • HW/SW Co-Design: Partner with the RTL and Architecture teams to change the custom ISA definitions.
Skills and Qualifications

Minimum Qualifications:

  • Bachelor’s degree in Electrical Engineering or Computer Engineering, and 5+ years practical industry experience working with advanced process nodes (7nm or below).
  • Deep hands-on experience with MLIR or XLA for deep learning workloads.
  • Expert-level understanding of PyTorch internals and how they interface with external backends.
  • Proficiency in modern C/C++ for building robust, scalable, and high-performance compiler infrastructure.
  • Advanced knowledge in Computer Architecture, especially the Programming Model, of at least one style of chip, including SoCs, CPUs, GPUs, or AI accelerators
  • A highly collaborative mindset to push boundaries and co-design effectively with other engineers.

Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

  • Hands-on experience in post-Silicon firmware and model update patches
  • Experience defining and implementing custom dialects, lowering passes, and graph rewrites in an LLVM-based ecosystem.
  • Knowledge of the tradeoffs between static and runtime environments, including JITs and ABIs
  • Location: This role is based in Austin, Texas or Palo Alto, California.
  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $420,000 USD, plus equity.
  • Visa Sponsorship: We sponsor visas. We can't guarantee success for every candidate or role, but if you're the right fit, we're committed to working through the visa process.
  • Benefits: River AI offers generous health, dental, and vision benefits, unlimited PTO, and relocation support as needed.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kernel Engineer (Custom Silicon), Hardware
Kernel Engineer (Custom Silicon), Hardware

River AI Inc. • Palo Alto (CA), Austin (TX)

On-site
USD 210,000 - 420,000
Health, dental, and vision benefits
Equity
Relocation support
Kernel Engineer (Custom Silicon), Hardware
Kernel Engineer (Custom Silicon), Hardware

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
Performance Engineer, Hardware
Performance Engineer, Hardware

River AI Inc. • Palo Alto (CA), Austin (TX)

On-site
USD 200,000 - 420,000
Health benefits
Unlimited PTO
Relocation assistance
Performance Engineer, Hardware
Performance Engineer, Hardware

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, vision benefits
Unlimited PTO
Relocation support
Compiler Engineer
Compiler Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Compiler Engineer for Custom Silicon & MLIR (Unlimited PTO)
AI Compiler Engineer for Custom Silicon & MLIR (Unlimited PTO)

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
Member of Technical Staff, AI-Driven Compilation
Member of Technical Staff, AI-Driven Compilation

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Meaningful equity
Front End Compiler
Front End Compiler

Lemurian Labs Inc. • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Equity
Medical/Dental/Vision
Retirement savings plan
+1
Member of Technical Staff, GPU Compiler
Member of Technical Staff, GPU Compiler

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Office in San Francisco
Member of Technical Staff, GPU Compiler
Member of Technical Staff, GPU Compiler

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance