Lead Kernel Engineer/Architect (m/f/d)

EPAM Systems

München

Hybrid

EUR 120.000 - 180.000

Vollzeit

Vor 5 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

30 days holiday per annum
Company Pension Scheme
Learning and development opportunities
Flexible and remote working options
Employee stock purchase program (ESPP)

Zusammenfassung

EPAM Systems is seeking a Lead Kernel Engineer/Architect in Germany to push the performance of AI workloads on GPUs/TPUs. You will design high-performance kernels, optimize ML operations for large-scale training and inference, and work with cloud platforms and OSS ecosystems.

The role requires collaboration with researchers, compiler engineers and framework developers. You will lead kernel optimizations, contribute to performance tooling, and drive architecture decisions across distributed

Qualifikationen

  • Bachelor’s degree or equivalent practical experience required.
  • 12+ years in software engineering or systems programming.
  • 5+ years of software development in C++ or Python.
  • 3+ years in testing, maintaining or launching software products; 1 year in architecture.
  • Hands-on kernel-level optimization for accelerators or HPC systems.

Aufgaben

  • Design and optimize high-performance kernels for TPU and GPU architectures.
  • Build and maintain performance infrastructure, benchmarking and autotuning systems.
  • Collaborate with ML framework and compiler teams to integrate custom kernels.
  • Track accelerator hardware and compiler advancements to improve performance.
  • Develop documentation, APIs and OSS components to improve usability.
  • Analyze and resolve complex performance bottlenecks in distributed systems.

Kenntnisse

12+ years software engineering
5+ years C++/Python
Kernel performance optimization
Architectural design experience
Cross-functional collaboration

Ausbildung

Bachelor’s degree or equivalent

Tools

CUDA
Triton
Pallas
MLIR/OpenXLA

Jobbeschreibung

We're looking for a Lead Kernel Engineer/Architect to join our team in Germany in a hybrid working mode.

Are you passionate about pushing advanced hardware accelerators to their limits? Join us in shaping the future of AI performance and scalability.

As a Lead Kernel Engineer/Architect, you will drive the optimization of critical machine learning operations for large-scale training and inference, working with cutting-edge hardware like TPUs and GPUs, advanced ML models and performance toolchains. Your work will enable faster AI research and production deployments on cloud platforms and within open-source ecosystems.

In this role, you will collaborate with researchers, compiler engineers and framework developers to deliver optimized, high-performance solutions that set the standard for modern AI computation.

Responsibilities
  • Design and optimize high-performance kernels for TPU and GPU architectures using low-level programming frameworks such as Pallas, Triton or Mosaic
  • Build and maintain performance infrastructure, including benchmarking suites, autotuning systems, regression testing frameworks and tooling
  • Collaborate with ML framework developers (e.g., JAX, PyTorch) and compiler teams (XLA/MLIR) to integrate custom kernels and reduce performance bottlenecks
  • Track advancements in accelerator hardware, compiler technology and AI model design to identify opportunities for kernel-level optimization
  • Develop clear documentation, APIs and supporting OSS components that improve developer usability and adoption
  • Analyze and resolve complex performance issues impacting large-scale distributed training and inference systems
Requirements
  • Bachelor’s degree or equivalent practical experience
  • 12+ years of industry experience in software engineering or systems programming
  • 5+ years of experience in software development using C++ or Python
  • 3+ years of experience in testing, maintaining or launching software products and at least 1 year in software design or architecture
  • Hands-on experience in performance optimization at the kernel level for accelerators or high-performance systems
Nice to have
  • Proficiency in low-level accelerator programming (CUDA, Triton, Pallas)
  • Familiarity with ML frameworks such as JAX or PyTorch and optimization techniques for attention layers, Mixture of Experts (MoE) and precision tuning
  • Strong understanding of modern hardware accelerators, including pipelining, data movement and heterogeneous compute
  • Knowledge of compiler principles and intermediate representations (e.g., MLIR, OpenXLA)
  • Experience building OSS developer infrastructure, APIs and performance-critical libraries
  • Excellent problem-solving skills and ability to collaborate in cross-functional engineering environments
We offer
  • 30 days holiday per annum
  • Company Pension Scheme
  • Regular performance assessments
  • Discount on Fitness-First Black Membership
  • Employee Stock Purchase Plan (ESPP) (subject to certain eligibility requirements)
  • Learning and development opportunities, including in-house training and coaching, professional certifications, and courses
  • Friendly and enjoyable working team
  • Regular corporate and social events
  • Flexible and remote working opportunities
  • Award-winning workplace: Great Place To Work certified in 2026, Kununu (Top Company 2022–2026), NewWork Business Award 2025 for outstanding culture, innovation and employee satisfaction.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Forward Deployed Engineer (m/f/d)
Senior Forward Deployed Engineer (m/f/d)

EPAM Systems • München

Hybrid
EUR 120.000 - 180.000
30 days holiday per annum
Company Pension Scheme
Regular performance assessments
+7
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

Remotestar • München

Hybrid
EUR 80.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
Senior Principal Computer Architect AI/ML (f/m/d)
Senior Principal Computer Architect AI/ML (f/m/d)

NXP Semiconductors • München

Vor Ort
EUR 80.000 - 110.000
Senior Principal Computer Architect AI/ML (f/m/d)
Senior Principal Computer Architect AI/ML (f/m/d)

NXP Semiconductors • Hamburg

Vor Ort
EUR 80.000 - 110.000
Strategic Technology Lead – AI Processors & Accelerators
Strategic Technology Lead – AI Processors & Accelerators

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • München

Vor Ort
EUR 120.000 - 180.000
AI Engineer (m/f/d)
AI Engineer (m/f/d)

KLA • Dresden

Hybrid
EUR 60.000 - 90.000
Team events
Fruits and drinks
Flexible working hours
Forward Deployed Engineer
Forward Deployed Engineer

turbalance • Heidelberg

Hybrid
EUR 60.000 - 80.000
Competitive compensation
Performance-based incentives
Subsidized Deutschlandticket
+2
AI Compiler Engineer (Senior Staff)
AI Compiler Engineer (Senior Staff)

roofline • Deutschland

Hybrid
EUR 90.000 - 120.000
Opportunities for career growth
Flexible work schedule
Equity options for employees
+1
Principal Engineer System Compute (f/m/div)
Principal Engineer System Compute (f/m/div)

Infineon Technologies • Frankfurt

Vor Ort
EUR 80.000 - 110.000
Linux System Administrator - R&D (m/f/d)
Linux System Administrator - R&D (m/f/d)

Stealth Physical AI • München

Vor Ort
EUR 70.000 - 110.000