Performance Engineer, Hardware

River AI Inc.

Palo Alto, Austin (CA, TX)

On-site

USD 200,000 - 420,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health benefits
Unlimited PTO
Relocation assistance

Job summary

River AI is seeking a hardware performance engineer to architect and model high-performance silicon, developing simulators to predict AI accelerator and SoC behavior for real-world models. You will own performance models, ISA/kernel optimization, and FPGA emulators for software development.

You will collaborate with compiler, IR, and RTL teams across the stack, validating models against RTL and emulators, and identifying bottlenecks from ILP to bandwidth limits within a fast-moving AI hardware

Qualifications

  • Bachelor’s degree in Electrical Engineering or Computer Engineering or Computer Science.
  • 5+ years of industry experience with advanced process nodes (7nm or below).
  • Proficiency in C/C++ or SystemC.
  • Experience with compilers (LLVM/GCC/XLA) and high-performance kernels (CUDA/Triton).
  • Expert knowledge in Computer Architecture of at least one chip style (SoCs/CPUs/GPUs/AI accelerators).
  • Experience profiling hardware with performance counters, hardware profilers, and trace analysis tools.

Responsibilities

  • Design and implement high-performance hardware models using C++ and/or SystemC.
  • Perform micro-architectural exploration to assess changes and impacts on IPC and execution time.
  • Profile AI kernels and software stacks to generate representative traces for stress tests.
  • Co-design with compiler and kernel teams to optimize software mapping to hardware.
  • Validate performance models against RTL and pre-silicon emulators for accuracy.
  • Identify and quantify system bottlenecks from ILP to bandwidth utilization.

Skills

C/C++
SystemC
LLVM/GCC/XLA
Computer Architecture
Performance profiling
Collaboration
RTL/Emulator

Education

Bachelor’s degree in Electrical/Computer Engineering or Computer Science

Tools

QEMU models
LLVM
GCC
CUDA/Triton

Job description

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are seeking exceptional hardware performance engineers to architect, model, and correlate high-performance custom silicon. You will develop high-fidelity simulators to predict how our AI accelerator architecture and SoC system will handle real-world AI models. You will take ownership of performance models, ISA and kernel optimization, and FPGA emulators for software development. You will be collaborating both up and down the stack with compiler, IR, and software teams, as well as with RTL design engineers.

What You’ll Do
  • Simulator Development: Design and implement high-performance and functional models of complex hardware using C++ and/or SystemC.
  • Micro-architectural Exploration: Conduct "what-if" studies to evaluate architectural changes (e.g., cache sizes, branch predictors, pipeline depths, scatter/gather, matmul shaping) and their impact on IPC, MFU, TTFT, and total execution time.
  • Workload Characterization: Analyze and profile AI kernels and software stacks to generate representative traces that stress-test the hardware models.
  • HW/SW Co-Design: Collaborate with compiler and kernel teams to optimize software mapping to hardware, ensuring the architecture supports emerging algorithmic breakthroughs efficiently.
  • Performance Correlation: Validate the performance model against RTL and pre-silicon emulators to ensure the model’s accuracy remains within strict tolerance levels.
  • Bottleneck Analysis: Identify and quantify system-level bottlenecks, ranging from instruction-level parallelism (ILP) limits to bandwidth throttling to utilization.
Skills and Qualifications

Minimum Qualifications:

  • Bachelor’s degree in Electrical Engineering or Computer Engineering or Computer Science, and 5+ years practical industry experience working with advanced process nodes (7nm or below).
  • Expert proficiency in C/C++ or event-driven simulation environments like SystemC
  • Hands-on experience with how compilers transform code (LLVM/GCC/XLA) and how high-performance kernels (CUDA/Triton) interact with the underlying ISA.
  • Expert knowledge in Computer Architecture of at least one style of chip, including SoCs, CPUs, GPUs, or AI accelerators
  • Experience with profiling hardware with performance counters, hardware profilers, and trace analysis tools to dissect application behavior.
  • A highly collaborative mindset to push boundaries and co-design effectively with other engineers.

Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

  • Hands-on experience in pre-silicon RTL/emulator and/or post-silicon performance validation
  • Knowledge or experience of QEMU models for pre-silicon software development
  • Proficiency in scripting for data post-processing, visualization of simulation results, and automation of massive regression suites.
  • Location: This role is based in Austin, Texas or Palo Alto, California.
  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $420,000 USD, plus equity.
  • Visa Sponsorship: We sponsor visas. We can't guarantee success for every candidate or role, but if you're the right fit, we're committed to working through the visa process.
  • Benefits: River AI offers generous health, dental, and vision benefits, unlimited PTO, and relocation support as needed.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RTL Design Engineer, Hardware
RTL Design Engineer, Hardware

River AI Inc. • Palo Alto (CA), Austin (TX)

On-site
USD 200,000 - 420,000
Health benefits
Dental benefits
Vision benefits
+2
Design Verification Engineer, Hardware
Design Verification Engineer, Hardware

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health benefits? (see listing)
Kernel Engineer (Custom Silicon), Hardware
Kernel Engineer (Custom Silicon), Hardware

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
Software Engineer, River API
Software Engineer, River API

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Performance Architect - AI Hardware
Performance Architect - AI Hardware

TEEMA • United States

Remote
USD 140,000 - 210,000
Physical Design Engineer, Hardware
Physical Design Engineer, Hardware

River AI Inc. • Palo Alto (CA), Austin (TX)

On-site
USD 200,000 - 420,000
Health, dental, vision benefits
Unlimited PTO
Relocation support
Performance Architect
Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment
Performance Modeling Engineer ~2
Performance Modeling Engineer ~2

OpenAI • Seattle (WA)

Hybrid
USD 266,000 - 445,000
Relocation assistance
Head of Performance Visibility
Head of Performance Visibility

Etched • San Jose (CA)

On-site
USD 200,000 - 300,000
Medical, dental, and vision packages with generous coverage
Housing subsidy of $2k per month
Relocation support
+1
Head of Performance Visibility
Head of Performance Visibility

Delos • San Jose (CA)

On-site
USD 150,000 - 200,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for new hires
+2