Principal Engineer – NPU Compiler & Architecture

Saur Energy International

West Virginia

On-site

USD 140,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Saur Energy International in the United States seeks a seasoned compiler architect to lead the development of compilers and software for next-generation NPUs. You will define software abstractions, lower graphs, map operators, and create lightweight runtime infrastructure to expose NPU capabilities to AI workloads through end-to-end prototype work.

As a member of the architecture team, you will collaborate with hardware and software colleagues, evaluate architectural trade-offs, and mentor

Qualifications

  • 8+ years of experience in compiler development, systems software, AI accelerators, or related areas.
  • Strong understanding of compiler architecture and optimization techniques.
  • Strong C++ programming skills and proficiency in Python.
  • Strong understanding of computer architecture and familiarity with accelerator architectures, memory hierarchies and data movement.
  • Experience analyzing AI/ML workloads and understanding performance bottlenecks.

Responsibilities

  • Develop compiler and software proof-of-concepts for current and next-generation NPU architectures.
  • Define software abstractions and compiler flows that expose NPU hardware capabilities to AI workloads.
  • Develop and evaluate compiler concepts including graph lowering, intermediate representations, operator mapping, scheduling, tiling, fusion, memory planning and code generation.
  • Translate NPU architectural concepts into executable software models and demonstrate feasibility using representative AI workloads.
  • Develop lightweight compiler/runtime infrastructure to validate new hardware features before production software.
  • Analyze existing compiler limitations and identify architectural changes to improve programmability and accelerator utilization.
  • Work as an integral member of the hardware architecture team to co-design NPU hardware and software.
  • Analyze how proposed hardware features can be exposed through the compiler and software stack.
  • Provide software-driven feedback on compute architecture, memory hierarchy, data movement, dataflow, scheduling, and accelerator programmability.
  • Identify hardware features that provide meaningful benefits to AI workloads and challenge features with insufficient software value.
  • Define compiler requirements and software abstractions for new NPU capabilities.
  • Participate in architecture reviews and influence hardware decisions from a software/workload perspective.
  • Evaluate architectural trade-offs considering performance, power, area, compiler complexity and software scalability.
  • Rapidly prototype software solutions for architectural concepts and demonstrate end-to-end execution of workloads on proposed NPU architectures.
  • Build proof-of-concepts for hardware architects to make informed decisions before RTL implementation.
  • Help establish software models and interfaces that can evolve into production compiler components.
  • Mentor engineers and contribute to technical direction for compiler-driven hardware/software co-design.

Skills

Compiler dev
C++ programming
Python
AI accelerators
Computer architecture
NPU architectures

Education

Bachelor's/Master's in CS/CE/EE

Tools

MLIR
LLVM
TOSA
Torch-MLIR

Job description

What You Will Be Responsible For
  • Compiler & Software Architecture for NPU
    • Develop compiler and software proof-of-concepts for current and next-generation NPU architectures.
    • Define software abstractions and compiler flows that efficiently expose NPU hardware capabilities to AI workloads.
    • Develop and evaluate compiler concepts including graph lowering, intermediate representations, operator mapping, scheduling, tiling, fusion, memory planning and code generation.
    • Translate NPU architectural concepts into executable software models and demonstrate their feasibility using representative AI workloads.
    • Develop lightweight compiler/runtime infrastructure to validate new hardware features before production software implementation.
    • Analyze existing compiler limitations and identify architectural changes required to improve programmability and accelerator utilization.
  • Hardware-Software Co-Design
    • Work as an integral member of the hardware architecture team to co-design NPU hardware and software.
    • Analyze how proposed hardware features can be effectively exposed through the compiler and software stack.
    • Provide software-driven feedback on compute architecture, memory hierarchy, data movement, dataflow, scheduling, instruction set and accelerator programmability.
    • Identify hardware features that provide meaningful benefits to real AI workloads and challenge features that add hardware complexity without sufficient software value.
    • Define compiler requirements and software abstractions for new NPU capabilities.
    • Participate in architecture and micro-architecture reviews and influence hardware decisions from a software and workload perspective.
    • Evaluate architectural trade-offs considering performance, power, area, compiler complexity and software scalability.
  • AI Workload & Performance Analysis
    • Analyze representative AI models and workloads to identify compute, memory, bandwidth, scheduling and data-movement bottlenecks.
    • Build software-based performance models and workload prototypes to evaluate architectural concepts.
    • Develop experiments to quantify the impact of proposed hardware features on model performance and accelerator utilization.
    • Investigate issues such as quantization, sparsity, operator fusion, tensor layouts, tiling, data reuse, memory bandwidth and scheduling efficiency.
    • Correlate software/model-level performance with architectural and micro-architectural behavior.
    • Use workload analysis to guide both current-generation optimizations and next-generation NPU architecture.
  • Architecture Prototyping
    • Rapidly prototype software solutions for architectural concepts that may be months or years away from production silicon.
    • Develop functional models, compiler prototypes, simulators, emulators, reference implementations or runtime abstractions as needed to validate architectural ideas.
    • Demonstrate end-to-end execution of representative AI workloads on proposed NPU architectures.
    • Build proof-of-concepts that allow hardware architects to make informed architectural decisions before RTL implementation.
    • Help establish software models and interfaces that can later evolve into production compiler components.
  • Cross-Functional Leadership
    • Work closely with NPU hardware architects, micro-architects, RTL designers and the production compiler/software organization.
    • Bridge the communication gap between hardware and software teams and translate requirements in both directions.
    • Participate in architecture definition from early concept through implementation and silicon bring-up.
    • Mentor engineers and contribute to technical direction for compiler-driven hardware/software co-design.
    • Influence the roadmap of future NPU architectures through workload and software-driven insights.
Necessary Qualifications

Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering or a related field.

8+ years of experience in compiler development, systems software, computer architecture, AI accelerators, or a closely related area.

Strong understanding of compiler architecture and compiler optimization techniques.

Strong C++ programming skills and proficiency in Python.

Strong understanding of computer architecture and familiarity with accelerator architectures, memory hierarchies and data movement.

Experience analyzing AI/ML workloads and understanding performance bottlenecks.

Ability to work across hardware and software boundaries and communicate effectively with hardware architects.

Strong analytical and problem-solving skills, with the ability to rapidly prototype and evaluate architectural ideas.

Good to have
  • Experience developing compilers or software stacks for NPUs, GPUs, DSPs, TPUs or other AI accelerators.
  • Experience with MLIR, LLVM, TOSA, StableHLO, Torch-MLIR or similar compiler infrastructure.
  • Experience with accelerator-specific scheduling, tiling, memory management or code generation.
  • Understanding of NPU architecture concepts such as dataflow, tensor engines, systolic/array-based compute, local SRAMs, DMA/data movement and accelerator instruction sets.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Tech Lead, AI Compiler
Tech Lead, AI Compiler

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 180,000 - 260,000
AI Compiler Engineer
AI Compiler Engineer

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 100,000 - 130,000
NPU Kernel/Operator Engineer
NPU Kernel/Operator Engineer

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 120,000 - 160,000
Compiler Engineer
Compiler Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior GPU Compiler Development Engineer
Senior GPU Compiler Development Engineer

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity opportunities
Comprehensive benefits package
Senior GPU Compiler Development Engineer
Senior GPU Compiler Development Engineer

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior GPU Compiler Development Engineer
Senior GPU Compiler Development Engineer

NVIDIA • Austin (TX)

On-site
USD 184,000 - 288,000
Stock options
Comprehensive benefits package
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

SF Tensor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Compiler Engineer
Compiler Engineer

Opticore • Berkeley (CA)

On-site
USD 120,000 - 180,000
Senior GPU Compiler Development Engineer
Senior GPU Compiler Development Engineer

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits