Compiler Expert

Adecco

Warszawa

On-site

PLN 280,000 - 420,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

On-site in Warsaw

Job summary

Adecco is seeking a highly capable engineer to join our deep learning infrastructure team in Warsaw, Poland. You will own the Triton compiler and kernel optimization framework for NPUs, driving AI workloads with advanced performance engineering and compiler techniques.

You will implement state-of-the-art Triton kernels, optimize memory patterns, and ensure tight integration with PyTorch and other runtimes. This role combines AI compilation, NPU programming, and system performance, with mentoring

Qualifications

  • Master’s or PhD in a relevant field and 5+ years in NPU/GPU programming or compiler development.
  • Strong knowledge of accelerator architectures and performance bottlenecks.
  • Proficiency in Triton, PTX, or LLVM IR for low-level optimization.

Responsibilities

  • Lead design and development of the Triton compiler and performance framework for NPUs.
  • Implement high-performance Triton kernels (Attention, MatMul, LayerNorm, Conv, Softmax).
  • Optimize memory access, scheduling, and cache usage for scalable performance.
  • Integrate Triton with PyTorch, XLA, and runtime stacks for end-to-end performance.
  • Apply auto-tuning, kernel fusion, and operator scheduling to maximize throughput.
  • Mentor teammates and establish processes for performance analysis and optimization.
  • Stay current with MLIR, TVM, and related compiler tech to push boundaries.
  • Perform workload fingerprinting for large models to guide optimizations.

Skills

NPU programming
Kernel optimization
Performance profiling
System design
Mentoring

Education

Master’s degree in Computer Architecture or related field

Tools

Triton
LLVM IR
PTX
TVM
MLIR
Cutlass

Job description

What can we offer:
  • On side job in Warsaw, Poland
  • Employment based on contract of employment (first probation period for 3 months)
  • Long term project
  • Work in one of the most advanced research development centers
About the Role:

You would join deep learning infrastructure team and take ownership of building a high-performance Triton compiler and kernel optimization framework, driving the next generation of AI workloads on NPUs. This is a highly technical role that sits at the intersection of AI compilation, NPU programming, and system performance engineering.

Key Responsibilities:
  • Lead the design and development of the Triton compiler and performance optimization framework, enabling high-performance operator implementations on NPUs.
  • Implement state-of-the-art Triton kernels (e.g., Attention, MatMul, LayerNorm, Conv, Softmax) with best-in-class efficiency.
  • Optimize memory access patterns and parallel scheduling, deeply understanding cache behavior, register allocation, and SM occupancy limits.
  • Drive end-to-end performance optimization by integrating Triton with framework backends (e.g., PyTorch, XLA) and runtime stacks..
  • Research and apply auto-tuning, kernel fusion, and operator scheduling technologies to maximize performance scalability.
  • Mentor team members in Triton kernel development and establish standard processes for performance analysis and optimization.
  • Stay on top of cutting-edge compiler technologies (MLIR, TVM, Hidet, Cutlass) and introduce innovative ideas to push performance boundaries.
  • Conduct performance modeling and workload fingerprinting for large models (LLM, Diffusion, etc.) to guide system-level optimization.
What We’re Looking For:
  • Master’s degree or above in Computer Architecture, Compiler Theory, High Performance Computing, or related field; PhD preferred.
  • 5+ years of experience in NPU/GPU programming, operator optimization, or compiler development.
  • Deep understanding of accelerator architectures and performance bottleneck analysis (compute units, vector lanes, memory hierarchy, caches, etc.)..
  • Proficiency in Triton, PTX, or LLVM IR for low-level programming and optimization.
  • Familiarity with PyTorch, TensorFlow, or JAX, and their graph execution and operator scheduling mechanisms.
  • Proven ability to independently develop, benchmark, and optimize complex kernels.
  • Skilled with performance profiling tools (e.g., perf, torch.profiler, and other vendor-neutral or runtime profilers) for quantitative analysis and performance modeling.
  • Strong system design and software engineering skills, balancing performance, maintainability, and generality.
Nice to have:
  • Open-source contributions to Triton, LLVM, TVM, MLIR, or PyTorch.
  • Experience with AI training or inference systems such as TensorRT, vLLM, DeepSpeed, OneFlow, or OpenXLA.
  • Publications or patents in kernel fusion, memory tiling, or async pipeline optimization.
  • Experience with distributed inference optimization (tensor/pipeline parallelism, ZeRO, PagedAttention).
  • Proven cross-platform optimization experience across different accelerator vendors and architectures.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Compiler Engineer
Compiler Engineer

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
Lead AI Compiler & Kernel Optimization for NPUs
Lead AI Compiler & Kernel Optimization for NPUs

microTECH Global LTD • Warszawa

On-site
PLN 260,000 - 380,000
Triton AI Compiler Architect for NPU Performance
Triton AI Compiler Architect for NPU Performance

Adecco • Warszawa

On-site
PLN 280,000 - 420,000
On-site in Warsaw
Senior GPU Compiler Software Development Engineer
Senior GPU Compiler Software Development Engineer

Luxoft • Kraków

On-site
PLN 120,000 - 180,000
Health & dental insurance
Internal mobility
More perks
Senior Deep Learning Compiler Engineer - PyTorch
Senior Deep Learning Compiler Engineer - PyTorch

NVIDIA • Warszawa

On-site
PLN 293,000 - 507,000
Acceleration Kernel Developer Lead
Acceleration Kernel Developer Lead

Tenstorrent • Warszawa

Hybrid
PLN 240,000 - 360,000
Highly competitive compensation package
Benefits included
Equal opportunity employer
Compiler Performance Engineer
Compiler Performance Engineer

AMD • Warszawa

On-site
PLN 100,000 - 130,000
Compiler Performance Engineer
Compiler Performance Engineer

Advanced Micro Devices • Warszawa

On-site
PLN 90,000 - 120,000
Senior GPU Compiler Engineer for OpenAI Triton on ROCm
Senior GPU Compiler Engineer for OpenAI Triton on ROCm

Luxoft • Kraków

On-site
PLN 120,000 - 180,000
Health & dental insurance
Internal mobility
More perks
C++ Machine Learning Engineer, Models Training
C++ Machine Learning Engineer, Models Training

Tenstorrent • Warszawa

Hybrid
PLN 180,000 - 280,000
Highly competitive compensation
Benefits package
Equal opportunity employer