GPU Performance Engineer

Ignite Next GmbH

München

Hybrid

EUR 85.000 - 110.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, versandbereit.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Ignite Next GmbH is seeking a GPU Performance Engineer to optimize demanding ML and HPC workloads. You will develop high‑performance CUDA kernels and work closely with compiler teams to turn profiling insights into automated optimizations.

The role requires strong C/C++/CUDA skills, deep GPU architecture knowledge, and hands-on profiling with Nsight tools. A Master’s degree and fluent English (German a plus) are expected.

Qualifikationen

  • Master's degree or equivalent in a technical field such as CS/EE/math.
  • Strong C/C++ and CUDA programming skills.
  • Deep understanding of modern GPU architectures and performance optimization.
  • Experience optimizing ML, HPC, or other GPU workloads.
  • Solid knowledge of GPU memory hierarchies, shared memory, caches, etc.
  • Experience profiling GPU apps with Nsight Compute/Systems or similar.
  • Strong Python skills for benchmarking and automation.
  • Proficient in Linux and Git-based workflows.
  • Excellent communication in English and/or German.

Aufgaben

  • Analyze and optimize performance-critical GPU workloads, esp. ML apps.
  • Benchmark, profile, and hand-tune GPU kernels across architectures.
  • Identify bottlenecks in memory, occupancy, and launch configurations.
  • Develop highly optimized CUDA kernels and GPU techniques.
  • Collaborate with compiler engineers to convert gains into compiler improvements.
  • Design reproducible benchmarks and performance evaluation pipelines.
  • Assess optimization quality across GPUs and vendors.
  • Contribute to performance models and heuristics.

Kenntnisse

C/C++ programming
CUDA programming
Python
GPU architecture
Profiling and benchmarking
Linux & Git workflows
English/German communication

Ausbildung

Master's degree in Computer Science, Electrical Engineering, Mathematics, or related field

Tools

Nsight Compute
Nsight Systems
CUPTI
rocProfiler
LLVM/MLIR
CUTLASS
Triton
cuBLAS
cuDNN

Jobbeschreibung

We make complex software run efficiently on any processor using a self-learning compiler and cloud-scale optimization infrastructure. Our team brings together researchers and engineers from RWTH Aachen, TU Munich, TU Darmstadt, and ETH Zurich to tackle some of the hardest problems in systems and infrastructure software. If you want to work on deeply technical challenges with real-world impact, join us and help shape the future of compute. As a GPU Performance Engineer at Daisytuner, you will operate at the intersection of GPU architecture, performance engineering, and compiler technology. You will analyze demanding machine learning and high-performance computing workloads, identify performance bottlenecks, and develop highly optimized GPU kernels. Working closely with our compiler engineers, you will translate performance insights into compiler transformations, optimization heuristics, and autotuning strategies that enable our compiler to automatically generate high-performance GPU code.

What you will be doing
  • Analyze and optimize performance-critical GPU workloads, with a strong focus on machine learning applications
  • Benchmark, profile, and hand-tune GPU kernels across modern accelerator architectures
  • Identify bottlenecks related to memory hierarchy, occupancy, instruction throughput, synchronization, and kernel launch configuration
  • Develop highly optimized CUDA kernels and GPU programming techniques for real-world workloads
  • Use profiling and performance analysis tools to understand kernel behavior and hardware utilization
  • Collaborate closely with compiler engineers to translate manual optimization techniques into automated compiler transformations and optimization strategies
  • Design reproducible benchmarking methodologies and performance evaluation pipelines
  • Evaluate optimization quality across different GPU architectures and vendors
  • Contribute to the development of performance models and heuristics for GPU optimization
  • Stay up to date with the latest GPU architectures, programming models, and optimization techniques
What you are bringing
  • Master's degree (or equivalent experience) in Computer Science, Electrical Engineering, Mathematics, or a related technical field
  • Strong C/C++ and CUDA programming skills
  • Deep understanding of modern GPU architectures and performance optimization
  • Experience optimizing machine learning, HPC, or other performance-critical GPU workloads
  • Solid understanding of GPU memory hierarchies, shared memory, caches, register usage, occupancy, warp scheduling, and instruction-level performance
  • Experience profiling GPU applications using tools such as NVIDIA Nsight Compute, Nsight Systems, CUPTI, rocProfiler, or similar
  • Ability to analyze low-level performance bottlenecks and systematically improve kernel efficiency
  • Strong Python skills for benchmarking, automation, and experimentation
  • Comfortable working in Linux environments and with Git-based workflows
  • Strong communication skills in English and/or German
  • A structured, analytical working style and willingness to take ownership in an early-stage environment
Nice to have
  • Experience with GPU programming frameworks such as CUTLASS, Triton, cuBLAS, cuDNN, ROCm, or SYCL
  • Familiarity with compiler infrastructures such as LLVM, MLIR, or similar systems
  • Experience with kernel fusion, code generation, or compiler optimizations
  • Knowledge of transformer architectures, LLM inference/training, or other modern ML workloads
  • Experience across multiple GPU vendors (NVIDIA, AMD, Intel)
What we offer
  • A small, highly technical team with direct impact on core technology
  • Competitive compensation and potential equity participation
  • The opportunity to work at the intersection of compilers, machine learning, and high-performance computing
  • Real ownership over GPU optimization strategies that directly influence the capabilities of our compiler and the performance of production workloads

We are an equal opportunity employer and welcome applications from people of all backgrounds. We value diversity and believe that different perspectives make us stronger. We do not discriminate based on gender, nationality, ethnic origin, religion, disability, age, sexual orientation, or identity.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Auto-tuning & Auto-scheduling Engineer
Auto-tuning & Auto-scheduling Engineer

Ignite Next GmbH • München

Hybrid
EUR 90.000 - 130.000
Equity opportunity
Competitive compensation
Compiler Engineer
Compiler Engineer

Ignite Next GmbH • München

Hybrid
EUR 90.000 - 135.000
Competitive compensation
Equity potential
Ownership over architecture
+2
Lead Kernel Engineer/Architect (m/f/d)
Lead Kernel Engineer/Architect (m/f/d)

EPAM Systems • München

Hybrid
EUR 140.000 - 200.000
30 days holiday
Company pension scheme
Regular performance assessments
+6
Lead Kernel Engineer/Architect (m/f/d)
Lead Kernel Engineer/Architect (m/f/d)

EPAM Systems • Berlin

Hybrid
EUR 120.000 - 180.000
30 days holiday per annum
Company Pension Scheme
Regular performance assessments
+1
Senior HPC Performance Engineer
Senior HPC Performance Engineer

NVIDIA • Deutschland

Vor Ort
EUR 100.000 - 140.000
Lead Kernel Engineer/Architect (m/f/d)
Lead Kernel Engineer/Architect (m/f/d)

EPAM Systems • Frankfurt

Hybrid
EUR 140.000 - 210.000
30 days holiday
Company Pension Scheme
Performance reviews
+5
CUDA Engineering Expert
CUDA Engineering Expert

aitrainer • Deutschland

Hybrid
EUR 50.000 - 80.000
Senior Developer Technology Engineer, CPU Performance
Senior Developer Technology Engineer, CPU Performance

NVIDIA • Deutschland

Hybrid
EUR 131.000 - 248.000
Equity
Benefits
Hybrid work model
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Obsidian • Berlin

Vor Ort
EUR 83.000 - 165.000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • Berlin

Vor Ort
EUR 60.000 - 90.000