Lead Kernel Engineer/Architect (m/f/d)

EPAM Systems

Germany (OH)

Hybrid

USD 104,919 - 151,550

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EPAM Systems is seeking a Lead Kernel Engineer/Architect to join our team in Germany, working in a hybrid model. This role involves designing high-performance kernels for TPU and GPU architectures, while pushing advanced hardware accelerators to their limits.

The ideal candidate will have extensive experience in software engineering and performance optimization, particularly at the kernel level. Interested professionals should apply to be a part of our mission to enhance AI performance and scalability.

Qualifications

  • 12+ years of industry experience in software engineering or systems programming.
  • 5+ years of experience in software development using C++ or Python.
  • Hands-on experience in performance optimization at the kernel level for accelerators.

Responsibilities

  • Design and optimize high-performance kernels for TPU and GPU architectures.
  • Build and maintain performance infrastructure including benchmarking and regression testing.
  • Collaborate with ML framework developers and compiler teams for kernel integration.

Skills

C++
Python
Performance optimization
Kernel programming

Education

Bachelor's degree or equivalent practical experience

Tools

CUDA
Triton
Pallas

Job description

Overview

We’re looking for a Lead Kernel Engineer/Architect to join our team in Germany in a hybrid working mode. Are you passionate about pushing advanced hardware accelerators to their limits? Join us in shaping the future of AI performance and scalability.

Responsibilities
  • Design and optimize high-performance kernels for TPU and GPU architectures using low-level programming frameworks such as Pallas, Triton or Mosaic.
  • Build and maintain performance infrastructure, including benchmarking suites, autotuning systems, regression testing frameworks and tooling.
  • Collaborate with ML framework developers (e.g., JAX, PyTorch) and compiler teams (XLA/MLIR) to integrate custom kernels and reduce performance bottlenecks.
  • Track advancements in accelerator hardware, compiler technology and AI model design to identify opportunities for kernel-level optimization.
  • Develop clear documentation, APIs and supporting OSS components that improve developer usability and adoption.
  • Analyze and resolve complex performance issues impacting large-scale distributed training and inference systems.
Qualifications
  • Bachelor’s degree or equivalent practical experience.
  • 12+ years of industry experience in software engineering or systems programming.
  • 5+ years of experience in software development using C++ or Python.
  • 3+ years of experience in testing, maintaining or launching software products and at least 1 year in software design or architecture.
  • Hands‑on experience in performance optimization at the kernel level for accelerators or high‑performance systems.
Nice to Have
  • Proficiency in low‑level accelerator programming (CUDA, Triton, Pallas).
  • Familiarity with ML frameworks such as JAX or PyTorch and optimization techniques for attention layers, Mixture of Experts (MoE) and precision tuning.
  • Strong understanding of modern hardware accelerators, including pipelining, data movement and heterogeneous compute.
  • Knowledge of compiler principles and intermediate representations (e.g., MLIR, OpenXLA).
  • Experience building OSS developer infrastructure, APIs and performance‑critical libraries.
  • Excellent problem‑solving skills and ability to collaborate in cross‑functional engineering environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Solution Architect (Kernel Optimization & ML Performance)
Solution Architect (Kernel Optimization & ML Performance)

EPAM Systems • Town of Poland (NY)

On-site
USD 140,000 - 180,000
Lead Kernel Architect for AI Accelerators
Lead Kernel Architect for AI Accelerators

EPAM Systems • Germany (OH)

Hybrid
USD 104,000 - 152,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
NPU Kernel/Operator Engineer
NPU Kernel/Operator Engineer

Black Sesame Technologies Inc • San Jose (CA)

On-site
USD 120,000 - 160,000
KERNEL ENGINEER
KERNEL ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Compiler Architect for LLM Accelerators
Senior ML Compiler Architect for LLM Accelerators

MRL Consulting Group Ltd. • Town of Texas (WI)

On-site
Lead Firmware Architect for AI Accelerator Hardware
Lead Firmware Architect for AI Accelerator Hardware

MRL Consulting Group Ltd. • Town of Texas (WI)

On-site
Founding GPU Kernel Engineer
Founding GPU Kernel Engineer

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 150,000 - 350,000
TPU Kernel Engineer for High-Performance ML Systems
TPU Kernel Engineer for High-Performance ML Systems

SignalAI • New York (NY)

Hybrid
USD 280,000 - 850,000