LPU Chip Architecture Engineer

Canaan Inc.

Singapore

On-site

SGD 120,000 - 190,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Canaan Inc. seeks a senior AI chip architect to lead LPU architecture, including static dataflow computing arrays and on-chip SRAM hierarchy.

You will collaborate with the compiler team to align hardware microarchitecture and drive hardware-software co-design, ensuring efficient inference performance. Responsibilities include performance benchmarking against GPU/NPU baselines, MoE/multimodal model tracking, FPGA prototype verification, and coordinating tape-out with back-end teams.

Qualifications

  • Master’s degree or above in Microelectronics, IC, or Computer Architecture with 2+ years in AI chip design.
  • Familiar with static dataflow and systolic architectures; knowledge of Prefill/Decode inference helpful.

Responsibilities

  • Lead architecture definition of LPU chips with static dataflow and SRAM hierarchy.
  • Collaborate with compiler team to define hardware microarchitecture and co-design.
  • Model power, bandwidth and latency; benchmark against GPU/NPU architectures.
  • Track MoE and multimodal model inferences; iterate architecture for next-gen inference.
  • Participate in chip front-end design, FPGA prototype verification, and tape-out with back-end teams.
  • Research dataflow chip architectures of overseas benchmarks and deliver competitive analyses.

Skills

AI chip architecture
Hardware-software co-design
Systolic array architecture
Static dataflow architecture
Performance benchmarking
Large model inference

Education

Master’s degree or above in Microelectronics/Integrated Circuit/Computer Architecture

Tools

Chip front-end design flow
Groq architecture familiarity

Job description

  • Lead the overall architecture definition of LPU chips based on the static dataflow architecture, including computing array design and on-chip SRAM storage hierarchy planning. Solve core pain points of high latency and frequent data movement in large model inference scenarios.
  • Collaborate with the compiler team in the early stage to define hardware microarchitecture and implement hardware-software co-design. Align scheduling logic in advance to avoid industry pain points caused by non-modifiable hardware scheduling after tape-out.
  • Conduct modeling and analysis on the computing power, bandwidth and power consumption of LPU chips. Carry out architecture performance benchmarking and optimization by comparing with traditional GPU and NPU architectures.
  • Track the inference requirements of MoE large models and multimodal large models, iterate and upgrade the LPU hardware architecture to adapt to next-generation large model inference scenarios.
  • Participate in chip front-end design and FPGA prototype verification, cooperate with the back-end team to complete chip tape-out, and follow up chip bring-up testing and performance optimization.
  • Research dataflow chip architectures of overseas benchmark manufacturers including Groq, Etched and Cerebras, and deliver competitive analysis reports and architecture iteration solutions.
  • Master’s degree or above in Microelectronics, Integrated Circuit, Computer Architecture or related majors, with no less than 2 years of working experience in AI chip architecture design.
  • Familiar with static dataflow architecture and systolic array architecture; understand the principles of Prefill and Decode dual-stage inference for large models. Prior experience in video memory and on-chip SRAM scheduling is preferred.
  • Hands-on experience in AI chip / NPU / GPU architecture design, familiar with chip front-end design flow, with solid hardware-software co-design capabilities.
  • Proficient in chip performance evaluation methodologies, capable of independently completing simulation and analysis of computing power, latency and power consumption.
  • R&D experience in dataflow chips or LPU chips.
  • Experience in hardware adaptation for large model inference chips.
  • Familiar with the fundamentals of AI compilers.
  • In-depth research experience on Groq architecture.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LPU Chip Architect: Dataflow AI Inference Expert
LPU Chip Architect: Dataflow AI Inference Expert

Canaan Inc. • Singapore

On-site
SGD 120,000 - 190,000
(Sr./Staff) NPU Design Engineer
(Sr./Staff) NPU Design Engineer

OMNIVISION TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 75,000 - 95,000
(Staff/Sr. Staff) NPU Design Engineer
(Staff/Sr. Staff) NPU Design Engineer

OMNIVISION • Singapore

On-site
SGD 90,000 - 130,000
System Performance Modeling Engineer/Architect (NPU)
System Performance Modeling Engineer/Architect (NPU)

Bitdeer Group • Singapore

On-site
SGD 80,000 - 120,000
Advanced Engineer (High-Efficiency AI Computing)
Advanced Engineer (High-Efficiency AI Computing)

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
(Staff/Sr. Staff) NPU Design Engineer
(Staff/Sr. Staff) NPU Design Engineer

OmniVision Technologies Singapore Pte. Ltd. • Singapore

On-site
SGD 120,000 - 180,000
Advanced Engineer (High-Efficiency AI Computing)
Advanced Engineer (High-Efficiency AI Computing)

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 180,000 - 300,000
Principal AI Chip Design Engineer(SG)
Principal AI Chip Design Engineer(SG)

Canaan Inc. • Singapore

On-site
SGD 180,000 - 260,000
Principal AI Chip Design Engineer
Principal AI Chip Design Engineer

CANAAN CREATIVE GLOBAL PTE. LTD. • Singapore

On-site
SGD 180,000 - 260,000
AI Computing Architecture Researcher
AI Computing Architecture Researcher

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000