Compiler Engineer

microTECH Global LTD

Warszawa

On-site

PLN 260,000 - 380,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

microTECH Global LTD is seeking a world-class compiler and performance optimization expert to join our deep learning infrastructure team in Warsaw. You will own the Triton compiler and kernel optimization framework, driving AI workloads on NPUs with advanced techniques in compiler theory and system performance engineering.

You will lead the design and implementation of state-of-the-art kernels, optimize memory patterns and scheduling, and integrate Triton with PyTorch and other backends to

Qualifications

  • Master's degree or above in Computer Architecture, Compiler Theory, High Performance Computing or related field.
  • PhD preferred for research track and advanced optimization tasks.
  • 5+ years of experience in NPU/GPU programming, operator optimization, or compiler development.

Responsibilities

  • Lead design and development of the Triton compiler and performance optimization framework for high-performance NPUs.
  • Implement state-of-the-art Triton kernels (Attention, MatMul, LayerNorm, Conv, Softmax) with top-tier efficiency.
  • Optimize memory access patterns, scheduling and cache usage to maximize SM occupancy and throughput.
  • Integrate Triton with PyTorch, XLA, and runtime stacks to enable end-to-end performance gains.
  • Research auto-tuning, kernel fusion, and operator scheduling to scale performance across models.
  • Mentor team members in Triton kernel development and establish standard performance-analysis processes.
  • Stay current with MLIR, TVM, Hidet, Cutlass and related compiler tech to push performance boundaries.
  • Perform performance modeling and workload fingerprinting for large models to guide optimizations.

Skills

NPU/GPU programming
Triton
Kernel development
Performance optimization
Auto-tuning
Kernel fusion
Operator scheduling
Memory optimization
Profiling tools

Education

Master's degree in Computer Architecture, Compiler Theory, HPC
PhD preferred

Tools

Triton
LLVM
TVM
MLIR
Cutlass
PyTorch

Job description

We are looking for a world-class compiler and performance optimization expert to join our deep learning infrastructure team. You will take ownership of building a high-performance Triton compiler and kernel optimization framework, driving the next generation of AI workloads on NPUs. This is a highly technical role that sits at the intersection of AI compilation, NPU programming,

and system performance engineering.

Key Responsibilities
  • Lead the design and development of the Triton compiler and performance optimization framework, enabling high-performance operator implementations on NPUs.
  • Implement state-of-the-art Triton kernels (e.g., Attention, MatMul, LayerNorm, Conv, Softmax) with best-in-class efficiency.
  • Optimize memory access patterns and parallel scheduling, deeply understanding cache behavior, register allocation, and SM occupancy limits.
  • Drive end-to-end performance optimization by integrating Triton with framework backends (e.g., PyTorch, XLA) and runtime stacks..
  • Research and apply auto-tuning, kernel fusion, and operator scheduling technologies to maximize performance scalability.
  • Mentor team members in Triton kernel development and establish standard processes for performance analysis and optimization.
  • Stay on top of cutting-edge compiler technologies (MLIR, TVM, Hidet, Cutlass) and introduce innovative ideas to push performance boundaries.
  • Conduct performance modeling and workload fingerprinting for large models (LLM, Diffusion, etc.) to guide system-level optimization.
What We're Looking For
  • Master's degree or above in Computer Architecture, Compiler Theory, High Performance Computing, or related field; PhD preferred.
  • 5+ years of experience in NPU/GPU programming, operator optimization, or compiler development.
  • Deep understanding of accelerator architectures and performance bottleneck analysis (compute units, vector lanes, memory hierarchy, caches, etc.)..
  • Proficiency in Triton, PTX, or LLVM IR for low-level programming and optimization.
  • Familiarity with PyTorch, TensorFlow, or JAX, and their graph execution and operator scheduling mechanisms.
  • Proven ability to independently develop, benchmark, and optimize complex kernels.
  • Skilled with performance profiling tools (e.g., perf, torch.profiler, and other vendor-neutral or runtime profilers) for quantitative analysis and performance modeling.
  • Strong system design and software engineering skills, balancing performance, maintainability, and generality.
Bonus Points
  • Open-source contributions to Triton, LLVM, TVM, MLIR, or PyTorch.
  • Experience with AI training or inference systems such as TensorRT, vLLM, DeepSpeed, OneFlow, or OpenXLA.
  • Publications or patents in kernel fusion, memory tiling, or async pipeline optimization.
  • Experience with distributed inference optimization (tensor/pipeline parallelism, ZeRO, PagedAttention).
  • Proven cross-platform optimization experience across different accelerator vendors and architectures.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Compiler & Kernel Optimization for NPUs
Lead AI Compiler & Kernel Optimization for NPUs

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
Senior GPU Compiler Software Development Engineer
Senior GPU Compiler Software Development Engineer

Luxoft • Kraków

On-site
PLN 120,000 - 180,000
Health & dental insurance
Internal mobility
More perks
Compiler Performance Engineer
Compiler Performance Engineer

AMD • Warszawa

On-site
PLN 100,000 - 130,000
Acceleration Kernel Developer Lead
Acceleration Kernel Developer Lead

Tenstorrent • Warszawa

Hybrid
PLN 240,000 - 360,000
Highly competitive compensation package
Benefits included
Equal opportunity employer
Senior GPU Compiler Engineer for OpenAI Triton on ROCm
Senior GPU Compiler Engineer for OpenAI Triton on ROCm

Luxoft • Kraków

On-site
PLN 120,000 - 180,000
Health & dental insurance
Internal mobility
More perks
Compiler Performance Engineer
Compiler Performance Engineer

Advanced Micro Devices • Warszawa

On-site
PLN 90,000 - 120,000
AI Researcher — Inference Optimization
AI Researcher — Inference Optimization

Featherless AI • Poland

Remote
PLN 60,000 - 80,000
Machine Learning Engineer — Inference Optimization
Machine Learning Engineer — Inference Optimization

Featherless AI • Poland

Remote
PLN 255,754 - 383,631
Competitive compensation
Meaningful equity at Series A
Real ownership over performance-critical systems
Senior C++ Software Engineer with CUDA/GPU/TPU
Senior C++ Software Engineer with CUDA/GPU/TPU

EPAM Systems • Poland

Hybrid
PLN 240,000 - 360,000
Hybrid by design
Remote work within Poland
Occasional international travel
+3
Solution Architect (Kernel Optimization & ML Performance)
Solution Architect (Kernel Optimization & ML Performance)

EPAM Systems • Warszawa

Hybrid
PLN 300,000 - 520,000
Hybrid work model within Poland
Opportunity to work abroad up to 60 d/
Relocation opportunities
+6