Compiler Engineer - Poland

microTECH Global Limited

Poland

Hybrid

PLN 180,000 - 280,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

microTECH Global Limited is seeking a world-class compiler and performance optimization expert to lead the Triton compiler and kernel framework for NPUs. You will drive AI workloads, implement advanced Triton kernels, and optimize memory, scheduling, and end-to-end performance with PyTorch and XLA.

The role demands a strong background in AI compilation, NPU programming, and systems performance engineering, with hands-on kernel development and mentoring responsibilities. PhD is a plus.

Qualifications

  • Master’s degree or above in Computer Architecture, Compiler Theory, High Performance Computing, or related field; PhD preferred.
  • 5+ years of experience in NPU/GPU programming, operator optimization, or compiler development.
  • Deep understanding of accelerator architectures and performance bottleneck analysis (compute units, vector lanes, memory hierarchy, caches).
  • Proficiency in Triton, PTX, or LLVM IR for low-level programming and optimization.
  • Familiarity with PyTorch, TensorFlow, or JAX, and their graph execution and operator scheduling mechanisms.
  • Proven ability to independently develop, benchmark, and optimize complex kernels.
  • Strong performance profiling and analysis skills with tools like perf, torch.profiler, etc.

Responsibilities

  • Lead design and development of the Triton compiler and performance optimization framework for high-performance operator implementations on NPUs.
  • Implement state-of-the-art Triton kernels (Attention, MatMul, LayerNorm, Conv, Softmax) with top-tier efficiency.
  • Optimize memory access patterns and parallel scheduling, understanding cache behavior, register allocation, and SM occupancy.
  • Drive end-to-end performance optimization by integrating Triton with PyTorch, XLA, and runtime stacks.
  • Research and apply auto-tuning, kernel fusion, and operator scheduling to maximize scalability.
  • Mentor team members in Triton kernel development and establish standard optimization processes.
  • Stay current with MLIR, TVM, and related compiler tech; push performance boundaries.
  • Conduct performance modeling and workload fingerprinting for large models to guide system optimization.

Skills

NPU programming
GPU programming
Triton
Operator optimization
LLVM IR
PyTorch
TensorFlow
Performance profiling
Performance modeling
Kernel development
System design

Education

Master’s degree or above in Computer Architecture

Tools

Triton
PTX
LLVM IR

Job description

We are looking for a world-classcompiler and performance optimization expert to join our deep learning infrastructure team. You will take ownership of building a high-performance Triton compiler and kernel optimization framework, driving the next generation of AI workloads on NPUs. This is a highly technical role that sits at the intersection ofAI compilation, NPU programming, and system performance engineering.

Key Responsibilities
  • Lead the design and development of the Triton compiler and performance optimization framework, enabling high-performance operator implementations on NPUs.
  • Implement state-of-the-art Triton kernels (e.g., Attention, MatMul, LayerNorm, Conv, Softmax) with best-in-class efficiency.
  • Optimize memory access patterns and parallel scheduling, deeply understanding cache behavior, register allocation, and SM occupancy limits.
  • Drive end-to-end performance optimization by integrating Triton with framework backends (e.g., PyTorch, XLA) and runtime stacks..
  • Research and apply auto-tuning, kernel fusion, and operator scheduling technologies to maximize performance scalability.
  • Mentor team members in Triton kernel development and establish standard processes for performance analysis and optimization.
  • Stay on top of cutting-edge compiler technologies (MLIR, TVM, Hidet, Cutlass) and introduce innovative ideas to push performance boundaries.
  • Conduct performance modeling and workload fingerprinting for large models (LLM, Diffusion, etc.) to guide system-level optimization.
What We’re Looking For
  • Master’s degree or above in Computer Architecture, Compiler Theory, High Performance Computing, or related field; PhD preferred.
  • 5+ years of experience in NPU/GPU programming, operator optimization, or compiler development.
  • Deep understanding of accelerator architectures and performance bottleneck analysis (compute units, vector lanes, memory hierarchy, caches, etc.)..
  • Proficiency in Triton, PTX, or LLVM IR for low-level programming and optimization.
  • Familiarity with PyTorch, TensorFlow, or JAX, and their graph execution and operator scheduling mechanisms.
  • Proven ability to independently develop, benchmark, and optimize complex kernels.
  • Skilled with performance profiling tools (e.g., perf, torch.profiler, and other vendor-neutral or runtime profilers) for quantitative analysis and performance modeling.
  • Strong system design and software engineering skills, balancing performance, maintainability, and generality.
Bonus Points
  • Open-source contributions to Triton, LLVM, TVM, MLIR, or PyTorch.
  • Experience with AI training or inference systems such as TensorRT, vLLM, DeepSpeed, OneFlow, or OpenXLA.
  • Publications or patents in kernel fusion, memory tiling, or async pipeline optimization.
  • Experience with distributed inference optimization (tensor/pipeline parallelism, ZeRO, PagedAttention).
  • Proven cross-platform optimization experience across different accelerator vendors and architectures.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Compiler Engineer
Compiler Engineer

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
Compiler Expert
Compiler Expert

Adecco • Warszawa

On-site
PLN 280,000 - 420,000
On-site in Warsaw
Triton Compiler & NPU Performance Engineer
Triton Compiler & NPU Performance Engineer

microTECH Global Limited • Poland

Hybrid
PLN 180,000 - 280,000
Lead AI Compiler & Kernel Optimization for NPUs
Lead AI Compiler & Kernel Optimization for NPUs

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
Triton AI Compiler Architect for NPU Performance
Triton AI Compiler Architect for NPU Performance

Adecco • Warszawa

On-site
PLN 280,000 - 420,000
On-site in Warsaw
Senior Deep Learning Compiler Engineer - PyTorch
Senior Deep Learning Compiler Engineer - PyTorch

NVIDIA • Warszawa

On-site
PLN 293,000 - 507,000
Acceleration Kernel Developer Lead
Acceleration Kernel Developer Lead

Tenstorrent • Warszawa

Hybrid
PLN 240,000 - 360,000
Highly competitive compensation package
Benefits included
Equal opportunity employer
Compiler Performance Engineer
Compiler Performance Engineer

AMD • Warszawa

On-site
PLN 100,000 - 130,000
Software Engineer - Triton Gdańsk, Pomeranian Voivodeship, Poland
Software Engineer - Triton Gdańsk, Pomeranian Voivodeship, Poland

Graphcore • Województwo pomorskie

Hybrid
PLN 120,000 - 180,000
Flexible working
Healthcare & dental cover
Phantom equity
+4
Compiler Performance Engineer
Compiler Performance Engineer

Advanced Micro Devices • Warszawa

On-site
PLN 90,000 - 120,000