Hardware Engineer (Lead Architect)

Normal Computing

United States

On-site

USD 180,000 - 260,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Normal Computing is seeking a Hardware Lead Architect to define the silicon and system microarchitecture for our unconventional AI accelerator, targeting orders of magnitude energy efficiency gains for LLM and diffusion workloads.

You will translate transformer workloads into compute tiles, memory hierarchies, and interconnects, and lead hardware/software co-design with RTL, compiler, and analog teams to maximize throughput per watt.

Qualifications

  • Experience with Python or C++ for performance modeling and analysis.
  • Proficiency with SystemVerilog/RTL and simulation-driven architecture.
  • Ability to read and translate workload profiles into datapath widths, pipelines, and area/power estimates.

Responsibilities

  • Define silicon and system microarchitecture for AI accelerator blocks and compute tiles.
  • Lead HW/SW co-design across RTL, compiler, and analog teams.
  • Build performance models and microarchitectural specifications for RTL implementation.
  • Prototyping and pre-silicon validation planning with RTL and FPGA teams.

Skills

Python
C++
SystemVerilog
RTL
Performance Modeling
Microarchitecture

Education

Electrical Engineering
Computer Engineering
Computer Science

Job description

  • As a Hardware Lead Architect, you will define the silicon and system microarchitecture for our custom unconventional compute platform—driving the architectural trade-offs that unlock a 100–1000x leap in energy efficiency over traditional digital chips for LLM and diffusion model inference
  • You will lead the hardware/software co-design efforts to break the von Neumann memory wall. By translating transformer architectures (KV-cache management, attention mechanisms) and diffusion execution flows into custom mixed-signal compute tiles, memory hierarchies, and tile interconnects, you will set the blueprint for our hardware
  • Working closely with compiler, RTL, and analog teams, you will build performance models, establish microarchitectural specifications, and ensure our custom silicon delivers maximum throughput-per-watt on real-world generative AI workloads
  • Compute Architecture: Help define the architecture and microarchitecture of novel AI accelerator compute blocks: PE array design, datapath organization, and support for efficiency techniques such as sparsity exploitation and reduced-precision computation. The compute tile is the surface where Normal’s research advantages have to show up in silicon, and you are one of the people responsible for making sure they do
  • Workload-to-Hardware Translation: Translate workload analysis and research findings into hardware specifications. Identify where architectural innovation creates the most leverage, define the structures that realize it, and produce microarchitecture documents unambiguous enough for RTL engineers to implement against. You work closely with them through implementation, not over the wall from it
  • Full-Stack PPA Tradeoffs: Reason across the full stack and defend PPA tradeoffs at every level. Move between algorithm-level workload behavior, memory hierarchy, on-chip interconnect, and physical design constraints. Make the call when the data is incomplete, and articulate why under scrutiny from our Systems Architect and the research team
  • ISA Co-Design: Partner with the compiler lead on ISA co-design. The programming model and the microarchitecture are defined together, and you are accountable for both sides meeting in the middle
  • Prototyping Strategy: Direct block-level pre-silicon validation. Decide which microarchitecture questions need to be answered, and the appropriate platform. Partner with our FPGA Design Engineers, who own implementation and bring-up, to de-risk decisions before tapeout. Work with the Systems Architect to make sure there are no gaps from block to System-level validation
  • Research Fluency: Stay current with the AI accelerator research landscape and be able to articulate clearly where Normal’s approach differs from existing solutions and why that matters. This is a research-adjacent seat and you are expected to read, possibly publish, and not just consume

Proficiency in Python or C++ for performance modeling and analysis, and familiarity with SystemVerilog or equivalent RTLExperience with simulation-driven architecture. You have used cycle-accurate or analytical models to make and defend design decisions before RTL exists, and you know which questions each tool can answer and which it cannotSubstantial experience in architecture or microarchitecture of high-performance digital systems: AI accelerators, compute engines, or similarly complex logic. You have shaped and directed the structures inside a chip, not just consumed them from the outsideExperience writing microarchitecture specifications and working closely with RTL engineers through implementationA degree in Electrical Engineering, Computer Engineering, Computer Science, or equivalent work experience. PhD welcome but not required; the bar is the work, not the credentialFluency moving between algorithm-level analysis and hardware specification. You can read a profile of a workload and translate it into datapath widths, pipeline stages, and area/power estimates without losing the thread on either sideComfort operating in an environment where the architecture is actively being discovered alongside the work. You do not need the answer to be already known to make progress on itFamiliarity with quantization and reduced-precision approaches for inference and their implementation implications. You understand the cost of a bit at the hardware level, not just the model level

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Engineer, Lead Architect
Hardware Engineer, Lead Architect

Drive Capital • Palo Alto (CA)

On-site
USD 260,000 - 360,000
Hardware Engineer, Architect
Hardware Engineer, Architect

Normal Computing • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Lead Hardware Architect: AI Accelerator & Efficient Compute
Lead Hardware Architect: AI Accelerator & Efficient Compute

Drive Capital • Palo Alto (CA)

On-site
USD 260,000 - 360,000
Research Engineer, Algorithms
Research Engineer, Algorithms

Normal Computing Corporation • New York (NY)

On-site
USD 140,000 - 210,000
Research Engineer, Algorithms
Research Engineer, Algorithms

Normal Computing • New York (NY)

On-site
USD 150,000 - 210,000
Lead AI Accelerator Hardware Architect
Lead AI Accelerator Hardware Architect

Normal Computing • United States

On-site
USD 180,000 - 260,000
Performance Architect - AI Hardware
Performance Architect - AI Hardware

TEEMA • United States

Remote
USD 140,000 - 210,000
AI Accelerator Hardware Architect
AI Accelerator Hardware Architect

Normal Computing • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Neuromorphic Front-End Lead
Neuromorphic Front-End Lead

Jobtailor • Arizona

On-site
USD 180,000 - 320,000
Research Scientist, Systems ML - HW/SW Co-Design
Research Scientist, Systems ML - HW/SW Co-Design

Meta • Menlo Park (CA)

On-site
USD 220,000 - 300,000