Triton Compiler & NPU Performance Engineer

microTECH Global Limited

Poland

Hybrid

PLN 180,000 - 280,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

microTECH Global Limited is seeking a world-class compiler and performance optimization expert to lead the Triton compiler and kernel framework for NPUs. You will drive AI workloads, implement advanced Triton kernels, and optimize memory, scheduling, and end-to-end performance with PyTorch and XLA.

The role demands a strong background in AI compilation, NPU programming, and systems performance engineering, with hands-on kernel development and mentoring responsibilities. PhD is a plus.

Qualifications

  • Master’s degree or above in Computer Architecture, Compiler Theory, High Performance Computing, or related field; PhD preferred.
  • 5+ years of experience in NPU/GPU programming, operator optimization, or compiler development.
  • Deep understanding of accelerator architectures and performance bottleneck analysis (compute units, vector lanes, memory hierarchy, caches).
  • Proficiency in Triton, PTX, or LLVM IR for low-level programming and optimization.
  • Familiarity with PyTorch, TensorFlow, or JAX, and their graph execution and operator scheduling mechanisms.
  • Proven ability to independently develop, benchmark, and optimize complex kernels.
  • Strong performance profiling and analysis skills with tools like perf, torch.profiler, etc.

Responsibilities

  • Lead design and development of the Triton compiler and performance optimization framework for high-performance operator implementations on NPUs.
  • Implement state-of-the-art Triton kernels (Attention, MatMul, LayerNorm, Conv, Softmax) with top-tier efficiency.
  • Optimize memory access patterns and parallel scheduling, understanding cache behavior, register allocation, and SM occupancy.
  • Drive end-to-end performance optimization by integrating Triton with PyTorch, XLA, and runtime stacks.
  • Research and apply auto-tuning, kernel fusion, and operator scheduling to maximize scalability.
  • Mentor team members in Triton kernel development and establish standard optimization processes.
  • Stay current with MLIR, TVM, and related compiler tech; push performance boundaries.
  • Conduct performance modeling and workload fingerprinting for large models to guide system optimization.

Skills

NPU programming
GPU programming
Triton
Operator optimization
LLVM IR
PyTorch
TensorFlow
Performance profiling
Performance modeling
Kernel development
System design

Education

Master’s degree or above in Computer Architecture

Tools

Triton
PTX
LLVM IR

Job description

microTECH Global Limited is seeking a world-class compiler and performance optimization expert to lead the Triton compiler and kernel framework for NPUs. You will drive AI workloads, implement advanced Triton kernels, and optimize memory, scheduling, and end-to-end performance with PyTorch and XLA.

The role demands a strong background in AI compilation, NPU programming, and systems performance engineering, with hands-on kernel development and mentoring responsibilities. PhD is a plus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Compiler & Kernel Optimization for NPUs
Lead AI Compiler & Kernel Optimization for NPUs

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
Compiler Engineer
Compiler Engineer

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
Compiler Engineer - Poland
Compiler Engineer - Poland

microTECH Global Limited • Poland

Hybrid
PLN 180,000 - 280,000
Compiler Expert
Compiler Expert

Adecco • Warszawa

On-site
PLN 280,000 - 420,000
On-site in Warsaw
Triton AI Compiler Architect for NPU Performance
Triton AI Compiler Architect for NPU Performance

Adecco • Warszawa

On-site
PLN 280,000 - 420,000
On-site in Warsaw
ML Frameworks & Performance Engineer
ML Frameworks & Performance Engineer

Graphcore • Województwo pomorskie

Hybrid
PLN 120,000 - 180,000
Flexible working
Healthcare & dental cover
Phantom equity
+4
ML Frameworks Engineer for AI Hardware (Triton/PyTorch)
ML Frameworks Engineer for AI Hardware (Triton/PyTorch)

Graphcore • Poland

Hybrid
PLN 150,000 - 210,000
Flexible working
Healthcare and dental
Phantom equity
+4
ML Frameworks Engineer (Triton) – AI Hardware & Software
ML Frameworks Engineer (Triton) – AI Hardware & Software

Graphcore • Województwo pomorskie

On-site
PLN 180,000 - 240,000
Flexible working
Generous leave
Pension matching
+4
Senior Deep Learning Compiler Engineer - PyTorch
Senior Deep Learning Compiler Engineer - PyTorch

NVIDIA • Warszawa

On-site
PLN 293,000 - 507,000
Software Engineer - Triton Gdańsk, Pomeranian Voivodeship, Poland
Software Engineer - Triton Gdańsk, Pomeranian Voivodeship, Poland

Graphcore • Województwo pomorskie

Hybrid
PLN 120,000 - 180,000
Flexible working
Healthcare & dental cover
Phantom equity
+4