NPU Research Engineer

Steneg

Berlin

Vor Ort

EUR 65.000 - 85.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

A leading global technology company in Berlin seeks an NPU Compiler & Framework Engineer to drive innovations in AI frameworks and compiler toolchains for NPU acceleration. This role involves the design of optimizations, collaboration with hardware teams, and contribution to the developer ecosystem. Candidates should possess a Master's degree in Computer Science and strong programming skills in Python and C/C++. Competitive compensation and growth opportunities are offered.

Qualifikationen

  • 3 years of experience in systems software or relevant PhD.
  • Solid knowledge of modern NPU/GPU architecture.
  • Hands-on experience with model optimization techniques.

Aufgaben

  • Design and implement optimizations for AI compilers targeting NPU workloads.
  • Automate model transformation and optimization on NPU hardware.
  • Collaborate with hardware teams to define architectural features.

Kenntnisse

Compiler internals (LLVM, GCC, TVM, XLA)
Programming in Python
Programming in C/C++
Communication skills in English

Ausbildung

Master’s degree in Computer Science, Computer Engineering, or related field
PhD in a relevant domain

Tools

AI frameworks (PyTorch, TensorFlow)
Compiler stacks (LLVM, TVM, GCC)

Jobbeschreibung

A leading global technology company driving innovation in AI acceleration and next-generation computing architectures. The organization develops high-performance hardware-software solutions focused on neural processing units (NPUs) and advanced compiler systems. Their products power large-scale AI systems across industries such as cloud computing, edge AI, and high-performance research infrastructure.

Mission

As an NPU Compiler & Framework Engineer, you will lead the co-design of AI frameworks and compiler toolchains tailored for NPU acceleration. Your work will directly impact how large AI models are executed with maximum efficiency, leveraging low-level optimizations, intelligent scheduling, and hardware-software synergy. This role is at the intersection of systems design, compiler architecture, and AI infrastructure, with strong visibility in both academic and open-source communities.

Responsibilities
NPU-Centric Framework and Runtime Design
  • Design and implement smart cross-layer optimizations for AI compilers and runtimes (e.g., PyTorch, vLLM) targeting NPU workloads.
  • Automate model transformation, quantization, and adaptive deployment for optimized execution on custom NPU hardware.
Compiler and Toolchain Development
  • Extend and optimize compiler stacks (LLVM, TVM, GCC) to translate high-level AI models into high-performance NPU code.
  • Focus on scheduling strategies, memory management, and parallel execution tailored to NPU microarchitectures.
Hardware-Software Co-Design
  • Collaborate with hardware teams to define ISA extensions, performance counters, and architectural features that enable better software-level optimization.
  • Participate in shaping the future of AI accelerators through feedback-driven development.
Research & Ecosystem Contribution
  • Publish results in top-tier systems and machine learning conferences (e.g., ISCA, ASPLOS, MLSys).
  • Support the developer ecosystem with documentation, tooling, and contributions to open-source AI compiler projects.
Required Qualifications
  • Master’s degree in Computer Science, Computer Engineering, or related field, plus 3 years of experience in systems software; or a recent PhD in a relevant domain.
  • Solid knowledge of compiler internals (LLVM, GCC, TVM, XLA) and modern NPU/GPU architecture.
  • Strong programming skills in Python and C/C++ for system-level development.
  • Excellent communication skills in English, both written and spoken.
Preferred Experience
  • Hands‑on experience with AI frameworks such as PyTorch or TensorFlow.
  • Familiarity with model optimization techniques, quantization, and graph transformation.
  • Experience working with DSP/xPU toolchains or specialized accelerators.
  • Open‑source contributions or publications in relevant conferences.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Compiler Engineer – AI Accelerators (LLVM/MLIR)
Senior Compiler Engineer – AI Accelerators (LLVM/MLIR)

AMD • Köln

Hybrid
EUR 90.000 - 130.000
AMD benefits
NPU Architect
NPU Architect

SEMRON GmbH • Dresden

Vor Ort
EUR 70.000 - 90.000
Senior Deep Learning Compiler Engineer - PyTorch
Senior Deep Learning Compiler Engineer - PyTorch

NVIDIA • Berlin

Vor Ort
EUR 110.000 - 160.000
Competitive salaries
Extensive benefits package
Diversity and inclusion
+1
Senior Deep Learning Compiler Engineer - PyTorch
Senior Deep Learning Compiler Engineer - PyTorch

NVIDIA • München

Vor Ort
EUR 75.000 - 100.000
Competitive salaries
Extensive benefits package
Diversity and inclusion initiatives
AI Compiler Engineer (Senior Staff)
AI Compiler Engineer (Senior Staff)

roofline • Köln

Vor Ort
EUR 80.000 - 110.000
Flexible working hours
Equity sharing
Team retreats
Senior Software Compute Architect (f/m/d)
Senior Software Compute Architect (f/m/d)

Black Semiconductor GmbH • Aachen

Hybrid
EUR 120.000 - 190.000
Senior Principal Computer Architect AI/ML (f/m/d)
Senior Principal Computer Architect AI/ML (f/m/d)

NXP Semiconductors • Hamburg

Vor Ort
EUR 80.000 - 110.000
Senior Principal Computer Architect AI/ML (f/m/d)
Senior Principal Computer Architect AI/ML (f/m/d)

NXP Semiconductors • München

Vor Ort
EUR 80.000 - 110.000
Senior Field Application Engineer – MPU, DNPU, Robotics, Drones
Senior Field Application Engineer – MPU, DNPU, Robotics, Drones

Jobtailor • München

Vor Ort
EUR 120.000 - 160.000
AI Compiler Engineer (Senior)
AI Compiler Engineer (Senior)

roofline • Köln

Vor Ort
EUR 70.000 - 90.000
Opportunities to grow
Flexible schedule
Equity participation
+1