Lead AI Compiler & Kernel Optimization for NPUs

microTECH Global LTD

Warszawa

On-site

PLN 260,000 - 380,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

microTECH Global LTD is seeking a world-class compiler and performance optimization expert to join our deep learning infrastructure team in Warsaw. You will own the Triton compiler and kernel optimization framework, driving AI workloads on NPUs with advanced techniques in compiler theory and system performance engineering.

You will lead the design and implementation of state-of-the-art kernels, optimize memory patterns and scheduling, and integrate Triton with PyTorch and other backends to

Qualifications

  • Master's degree or above in Computer Architecture, Compiler Theory, High Performance Computing or related field.
  • PhD preferred for research track and advanced optimization tasks.
  • 5+ years of experience in NPU/GPU programming, operator optimization, or compiler development.

Responsibilities

  • Lead design and development of the Triton compiler and performance optimization framework for high-performance NPUs.
  • Implement state-of-the-art Triton kernels (Attention, MatMul, LayerNorm, Conv, Softmax) with top-tier efficiency.
  • Optimize memory access patterns, scheduling and cache usage to maximize SM occupancy and throughput.
  • Integrate Triton with PyTorch, XLA, and runtime stacks to enable end-to-end performance gains.
  • Research auto-tuning, kernel fusion, and operator scheduling to scale performance across models.
  • Mentor team members in Triton kernel development and establish standard performance-analysis processes.
  • Stay current with MLIR, TVM, Hidet, Cutlass and related compiler tech to push performance boundaries.
  • Perform performance modeling and workload fingerprinting for large models to guide optimizations.

Skills

NPU/GPU programming
Triton
Kernel development
Performance optimization
Auto-tuning
Kernel fusion
Operator scheduling
Memory optimization
Profiling tools

Education

Master's degree in Computer Architecture, Compiler Theory, HPC
PhD preferred

Tools

Triton
LLVM
TVM
MLIR
Cutlass
PyTorch

Job description

microTECH Global LTD is seeking a world-class compiler and performance optimization expert to join our deep learning infrastructure team in Warsaw. You will own the Triton compiler and kernel optimization framework, driving AI workloads on NPUs with advanced techniques in compiler theory and system performance engineering.

You will lead the design and implementation of state-of-the-art kernels, optimize memory patterns and scheduling, and integrate Triton with PyTorch and other backends to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Compiler Engineer
Compiler Engineer

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
Senior C++ ML/AI Kernel & Compiler Engineer
Senior C++ ML/AI Kernel & Compiler Engineer

EPAM Systems • Poland

Hybrid
PLN 240,000 - 360,000
Hybrid by design
Remote work within Poland
Occasional international travel
+3
Senior PyTorch Deep Learning Compiler Engineer
Senior PyTorch Deep Learning Compiler Engineer

NVIDIA Corporation • Warszawa

On-site
PLN 292,000 - 507,000
Extensive benefits package
Flexible work environment
Diversity and inclusion initiatives
Senior Solution Architect: Kernel & ML Performance
Senior Solution Architect: Kernel & ML Performance

EPAM Systems • Kraków

Hybrid
PLN 400,000 - 640,000
Hybrid by design
Remote work within Poland
Relocation opportunities
+1
ML Performance Architect: Kernel & Systems Optimizations
ML Performance Architect: Kernel & Systems Optimizations

EPAM Systems • Warszawa

Hybrid
PLN 300,000 - 520,000
Hybrid work model within Poland
Opportunity to work abroad up to 60 d/
Relocation opportunities
+6
Senior GPU Compiler Software Development Engineer
Senior GPU Compiler Software Development Engineer

Luxoft • Kraków

On-site
PLN 120,000 - 180,000
Health & dental insurance
Internal mobility
More perks
Acceleration Kernel Developer Lead
Acceleration Kernel Developer Lead

Tenstorrent • Warszawa

Hybrid
PLN 240,000 - 360,000
Highly competitive compensation package
Benefits included
Equal opportunity employer
Senior C++ Software Engineer with CUDA/GPU/TPU
Senior C++ Software Engineer with CUDA/GPU/TPU

EPAM Systems • Poland

Hybrid
PLN 240,000 - 360,000
Hybrid by design
Remote work within Poland
Occasional international travel
+3
Senior Compiler Performance Engineer: Optimize AI/CPU
Senior Compiler Performance Engineer: Optimize AI/CPU

AMD • Warszawa

On-site
PLN 100,000 - 130,000
Senior GPU Compiler Engineer for OpenAI Triton on ROCm
Senior GPU Compiler Engineer for OpenAI Triton on ROCm

Luxoft • Kraków

On-site
PLN 120,000 - 180,000
Health & dental insurance
Internal mobility
More perks