Principal Engineer - NPU Compiler & Architecture

NXP Semiconductors

Hyderabad

On-site

INR 4,000,000 - 6,000,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NXP Semiconductors in Hyderabad seeks a seasoned Compiler/Software Architect to co-design NPU hardware and software. You will create compiler proofs-of-concept, define flows, and translate architectural ideas into executable models while evaluating AI workloads for performance and efficiency.

The role requires deep expertise in graph lowering, IRs, operator mapping, scheduling, tiling and code generation, plus collaboration with hardware teams to push forward production-ready software stacks.

Qualifications

  • Experience in compiler and software stack development for AI accelerators.
  • Ability to translate architectural concepts into executable software models.
  • Strong collaboration with hardware and software teams.

Responsibilities

  • Develop compiler and software proof-of-concepts for current and next-gen NPU architectures.
  • Define software abstractions and compiler flows exposing NPU capabilities to AI workloads.
  • Develop and evaluate compiler concepts including graph lowering, IR, operator mapping, scheduling, tiling, fusion, memory planning and code generation.
  • Translate NPU concepts into executable software models and validate feasibility against AI workloads.
  • Develop lightweight compiler/runtime infra to validate new hardware features before production.
  • Analyze compiler limitations and propose architectural changes for better programmability.
  • Co-design NPU hardware and software with architecture teams and RTL designers.
  • Provide software-driven feedback on compute architecture, memory, data movement and dataflow.
  • Define compiler requirements and abstractions for new NPU capabilities.

Skills

Compiler
Software architecture
Graph lowering
Intermediate representations
Operator mapping
Scheduling
Tiling
Fusion
Memory planning
Code generation
Performance modeling
Prototyping

Job description

About the Role

We are seeking a highly experienced Compiler / Software Architect to join our NPU Hardware Architecture team and play a key role in the hardware-software co-design of next-generation AI inference accelerators.

This is a unique role at the intersection of NPU hardware architecture, AI compilers, performance modeling, and software development. The position will work directly with hardware architects to define, evaluate, and prototype the software stack required to program current and future NPU architectures.

Our NPU architecture is closely coupled with the compiler stack. Efficient utilization of the accelerator requires deep understanding of the hardware compute architecture, memory hierarchy, dataflow, scheduling, quantization, instruction set and execution model. The successful candidate will use this understanding to develop compiler and software proof-of-concepts, evaluate architectural proposals against real AI workloads, and influence the design of future NPU hardware.

You will work closely with hardware architects, micro‑architects, compiler engineers and AI software teams to answer a fundamental question:

How should the NPU hardware and software be co‑designed to deliver the best performance, power efficiency, programmability and scalability for real‑world AI workloads?

What You Will Be Responsible For
Compiler & Software Architecture for NPU
  • Develop compiler and software proof-of-concepts for current and next‑generation NPU architectures.
  • Define software abstractions and compiler flows that efficiently expose NPU hardware capabilities to AI workloads.
  • Develop and evaluate compiler concepts including graph lowering, intermediate representations, operator mapping, scheduling, tiling, fusion, memory planning and code generation.
  • Translate NPU architectural concepts into executable software models and demonstrate their feasibility using representative AI workloads.
  • Develop lightweight compiler/runtime infrastructure to validate new hardware features before production software implementation.
  • Analyze existing compiler limitations and identify architectural changes required to improve programmability and accelerator utilization.
Hardware-Software Co-Design
  • Work as an integral member of the hardware architecture team to co‑design NPU hardware and software.
  • Analyze how proposed hardware features can be effectively exposed through the compiler and software stack.
  • Provide software‑driven feedback on compute architecture, memory hierarchy, data movement, dataflow, scheduling, instruction set and accelerator programmability.
  • Identify hardware features that provide meaningful benefits to real AI workloads and challenge features that add hardware complexity without sufficient software value.
  • Define compiler requirements and software abstractions for new NPU capabilities.
  • Participate in architecture and micro‑architecture reviews and influence hardware decisions from a software and workload perspective.
AI Workload & Performance Analysis
  • Analyze representative AI models and workloads to identify compute, memory, bandwidth, scheduling and data‑movement bottlenecks.
  • Build software‑based performance models and workload prototypes to evaluate architectural concepts.
  • Develop experiments to quantify the impact of proposed hardware features on model performance and accelerator utilization.
  • Investigate issues such as quantization, sparsity, operator fusion, tensor layouts, tiling, data reuse, memory bandwidth and scheduling efficiency.
  • Correlate software/model‑level performance with architectural and micro‑architectural behavior.
  • Use workload analysis to guide both current‑generation optimizations and next‑generation NPU architecture.
Architecture Prototyping
  • Rapidly prototype software solutions for architectural concepts that may be months or years away from production silicon.
  • Develop functional models, compiler prototypes, simulators, emulators, reference implementations or runtime abstractions as needed to validate architectural ideas.
  • Demonstrate end‑to‑end execution of representative AI workloads on proposed NPU architectures.
  • Build proof‑of‑concepts that allow hardware architects to make informed architectural decisions before RTL implementation.
  • Help establish software models and interfaces that can later evolve into production compiler components.
Cross-Functional Leadership
  • Work closely with NPU hardware architects, micro‑architects, RTL designers and the production compiler/software organization.
  • Bridge the communication gap between hardware and software teams and translate requirements in both directions.
  • Participate in architecture definition from early concept through implementation and silicon bring‑up.
  • Mentor engineers and contribute to technical direction for compiler‑driven hardware/software co‑design.
  • Influence the roadmap of future NPU architectures through workload
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Engineer – NPU Compiler & Architecture
Principal Engineer – NPU Compiler & Architecture

NXP Semiconductors • Hyderabad

On-site
INR 4,000,000 - 7,000,000
AI Accelerator Chip Architect
AI Accelerator Chip Architect

Capgemini • Bengaluru

On-site
INR 3,500,000 - 5,200,000
Senior Hardware ASIC Arch/Design Engineer
Senior Hardware ASIC Arch/Design Engineer

NXP Semiconductors • Hyderabad

On-site
INR 1,500,000 - 2,000,000
Lead Software Engineer - AI Operators
Lead Software Engineer - AI Operators

Kinara.ai • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Senior Applied Research Engineer - Accelerator Programming Model and Compiler
Senior Applied Research Engineer - Accelerator Programming Model and Compiler

NVIDIA Gruppe • Bengaluru

On-site
INR 6,000,000 - 9,000,000
NPU/AI Processor Synthesis - Principal Engineer/Manager
NPU/AI Processor Synthesis - Principal Engineer/Manager

Qualcomm • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Server Performance Architect - Hardware
Server Performance Architect - Hardware

NVIDIA Gruppe • Bengaluru

On-site
INR 4,000,000 - 9,000,000
Manager, Compiler Engineering - GPU
Manager, Compiler Engineering - GPU

NVIDIA Gruppe • Bengaluru

On-site
INR 4,500,000 - 8,000,000
GPU Architect
GPU Architect

NVIDIA AI • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Sr. Software Engineer – AI Operators
Sr. Software Engineer – AI Operators

NXP Semiconductors • Hyderabad

On-site
INR 2,500,000 - 3,500,000