Principal Engineer – NPU Compiler & Architecture

Saur Energy International

West Virginia

On-site

USD 140,000 - 200,000

Full time

5 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Saur Energy International is seeking a senior compiler and software architect to drive the design of NPU compiler flows and software abstractions for current and future architectures. You will collaborate with hardware teams to expose NPU capabilities to AI workloads and prototype software models to validate architectural ideas.

The role emphasizes cross-functional leadership, rapid prototyping, and evaluating trade-offs across performance, power, and data movement while shaping next-generation

Qualifications

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering or a related field.
  • 8+ years of experience in compiler development, systems software, computer architecture, AI accelerators, or a closely related area.
  • Strong understanding of compiler architecture and compiler optimization techniques.
  • Strong C++ programming skills and proficiency in Python.
  • Strong understanding of computer architecture and familiarity with accelerator architectures, memory hierarchies and data movement.
  • Experience analyzing AI/ML workloads and understanding performance bottlenecks.
  • Ability to work across hardware and software boundaries and communicate effectively with hardware architects.
  • Strong analytical and problem-solving skills, with the ability to rapidly prototype and evaluate architectural ideas.

Responsibilities

  • Develop compiler and software proof-of-concepts for current and next-generation NPU architectures.
  • Define software abstractions and compiler flows that efficiently expose NPU hardware capabilities to AI workloads.
  • Develop and evaluate compiler concepts including graph lowering, intermediate representations, operator mapping, scheduling, tiling, fusion, memory planning and code generation.
  • Translate NPU architectural concepts into executable software models and demonstrate their feasibility using representative AI workloads.
  • Develop lightweight compiler/runtime infrastructure to validate new hardware features before production software implementation.
  • Analyze existing compiler limitations and identify architectural changes required to improve programmability and accelerator utilization.
  • Work as an integral member of the hardware architecture team to co-design NPU hardware and software.
  • Analyze how proposed hardware features can be effectively exposed through the compiler and software stack.
  • Provide software-driven feedback on compute architecture, memory hierarchy, data movement, dataflow, scheduling, instruction set and accelerator programmability.
  • Identify hardware features that provide meaningful benefits to real AI workloads and challenge features that add hardware complexity without sufficient software value.
  • Define compiler requirements and software abstractions for new NPU capabilities.
  • Participate in architecture and micro-architecture reviews and influence hardware decisions from a software and workload perspective.
  • Evaluate architectural trade-offs considering performance, power, area, compiler complexity and software scalability.
  • Analyze representative AI models and workloads to identify compute, memory, bandwidth, scheduling and data-movement bottlenecks.
  • Build software-based performance models and workload prototypes to evaluate architectural concepts.
  • Develop experiments to quantify the impact of proposed hardware features on model performance and accelerator utilization.
  • Investigate issues such as quantization, sparsity, operator fusion, tensor layouts, tiling, data reuse, memory bandwidth and scheduling efficiency.
  • Correlate software/model-level performance with architectural and micro-architectural behavior.
  • Use workload analysis to guide both current-generation optimizations and next-generation NPU architecture.
  • Rapidly prototype software solutions for architectural concepts that may be months or years away from production silicon.
  • Develop functional models, compiler prototypes, simulators, emulators, reference implementations or runtime abstractions as needed to validate architectural ideas.
  • Demonstrate end-to-end execution of representative AI workloads on proposed NPU architectures.
  • Build proof-of-concepts that allow hardware architects to make informed architectural decisions before RTL implementation.
  • Help establish software models and interfaces that can later evolve into production compiler components.
  • Mentor engineers and contribute to technical direction for compiler-driven hardware/software co-design.
  • Influence the roadmap of future NPU architectures through workload and software-driven insights.

Skills

C++ programming
Python
Compiler development
AI accelerators
Computer architecture
Performance analysis

Education

Bachelor's or Master's in CS/CE/EE

Tools

MLIR
LLVM
TOSA
Torch-MLIR
StableHLO

Job description

Compiler & Software Architecture for NPU
  • Develop compiler and software proof-of-concepts for current and next-generation NPU architectures.
  • Define software abstractions and compiler flows that efficiently expose NPU hardware capabilities to AI workloads.
  • Develop and evaluate compiler concepts including graph lowering, intermediate representations, operator mapping, scheduling, tiling, fusion, memory planning and code generation.
  • Translate NPU architectural concepts into executable software models and demonstrate their feasibility using representative AI workloads.
  • Develop lightweight compiler/runtime infrastructure to validate new hardware features before production software implementation.
  • Analyze existing compiler limitations and identify architectural changes required to improve programmability and accelerator utilization.
Hardware–Software Co-Design
  • Work as an integral member of the hardware architecture team to co-design NPU hardware and software.
  • Analyze how proposed hardware features can be effectively exposed through the compiler and software stack.
  • Provide software-driven feedback on compute architecture, memory hierarchy, data movement, dataflow, scheduling, instruction set and accelerator programmability.
  • Identify hardware features that provide meaningful benefits to real AI workloads and challenge features that add hardware complexity without sufficient software value.
  • Define compiler requirements and software abstractions for new NPU capabilities.
  • Participate in architecture and micro-architecture reviews and influence hardware decisions from a software and workload perspective.
  • Evaluate architectural trade-offs considering performance, power, area, compiler complexity and software scalability.
AI Workload & Performance Analysis
  • Analyze representative AI models and workloads to identify compute, memory, bandwidth, scheduling and data-movement bottlenecks.
  • Build software-based performance models and workload prototypes to evaluate architectural concepts.
  • Develop experiments to quantify the impact of proposed hardware features on model performance and accelerator utilization.
  • Investigate issues such as quantization, sparsity, operator fusion, tensor layouts, tiling, data reuse, memory bandwidth and scheduling efficiency.
  • Correlate software/model-level performance with architectural and micro-architectural behavior.
  • Use workload analysis to guide both current-generation optimizations and next-generation NPU architecture.
Architecture Prototyping
  • Rapidly prototype software solutions for architectural concepts that may be months or years away from production silicon.
  • Develop functional models, compiler prototypes, simulators, emulators, reference implementations or runtime abstractions as needed to validate architectural ideas.
  • Demonstrate end-to-end execution of representative AI workloads on proposed NPU architectures.
  • Build proof-of-concepts that allow hardware architects to make informed architectural decisions before RTL implementation.
  • Help establish software models and interfaces that can later evolve into production compiler components.
Cross-Functional Leadership
  • Work closely with NPU hardware architects, micro-architects, RTL designers and the production compiler/software organization.
  • Bridge the communication gap between hardware and software teams and translate requirements in both directions.
  • Participate in architecture definition from early concept through implementation and silicon bring-up.
  • Mentor engineers and contribute to technical direction for compiler-driven hardware/software co-design.
  • Influence the roadmap of future NPU architectures through workload and software-driven insights.
Necessary Qualifications

Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering or a related field.

8+ years of experience in compiler development, systems software, computer architecture, AI accelerators, or a closely related area.

Strong understanding of compiler architecture and compiler optimization techniques.

Strong C++ programming skills and proficiency in Python.

Strong understanding of computer architecture and familiarity with accelerator architectures, memory hierarchies and data movement.

Experience analyzing AI/ML workloads and understanding performance bottlenecks.

Ability to work across hardware and software boundaries and communicate effectively with hardware architects.

Strong analytical and problem-solving skills, with the ability to rapidly prototype and evaluate architectural ideas.

Good to have
  • Experience developing compilers or software stacks for NPUs, GPUs, DSPs, TPUs or other AI accelerators.
  • Experience with MLIR, LLVM, TOSA, StableHLO, Torch-MLIR or similar compiler infrastructure.
  • Experience with accelerator-specific scheduling, tiling, memory management or code generation.
  • Understanding of NPU architecture concepts such as dataflow, tensor engines, systolic/array-based compute, local SRAMs, DMA/data movement and accelerator instruction sets.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Tech Lead, AI Compiler
Tech Lead, AI Compiler

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 180,000 - 260,000
NPU Kernel/Operator Engineer
NPU Kernel/Operator Engineer

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 120,000 - 160,000
AI Compiler Engineer
AI Compiler Engineer

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 100,000 - 130,000
NPU Compiler Architect – AI Hardware/Software Co-Design
NPU Compiler Architect – AI Hardware/Software Co-Design

Saur Energy International • West Virginia

On-site
USD 140,000 - 200,000
AI Compiler Tech Lead
AI Compiler Tech Lead

Darwin Recruitment • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior GPU Compiler Development Engineer
Senior GPU Compiler Development Engineer

NVIDIA • Austin (TX)

On-site
USD 184,000 - 287,500
Stock options
Comprehensive benefits package
Senior GPU Compiler Development Engineer
Senior GPU Compiler Development Engineer

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 287,500
Equity opportunities
Comprehensive benefits package
Principal AI Compiler Engineer
Principal AI Compiler Engineer

NXP Semiconductors NV • Austin (TX)

On-site
USD 180,000 - 240,000
Principal Compiler Engineer
Principal Compiler Engineer

Oho Group • United States

On-site
USD 180,000 - 280,000
Principal Compiler Engineer
Principal Compiler Engineer

Mulya Consulting • United States

Hybrid
USD 26,000 - 54,000