AI Research Engineer (Kernel & Inference Optimization) arbeitnow Jobgether Germany · 9/30/2026

Primetime

Deutschland

Remote

EUR 110.000 - 160.000

Vollzeit

Vor 6 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Remote-first working environment

Zusammenfassung

Primetime is seeking an AI Research Engineer (Kernel & Inference Optimization) based in Germany. You will work at the intersection of AI research and systems engineering to develop and optimize model-serving architectures for diverse hardware, including mobile and edge devices.

You will tackle latency, throughput, and memory efficiency challenges, designing custom GPU kernels and GPU shaders in Metal Shading Language.

Qualifikationen

  • PhD in NLP/ML or related with a strong AI research track and publications.
  • Proven expertise in Metal Shading Language (MSL) and custom compute shaders.

Aufgaben

  • Design and deploy high-throughput, low-latency model-serving architectures.
  • Develop inference pipelines for mobile and edge devices.
  • Establish performance targets for latency, throughput, memory, and reliability.
  • Build controlled inference benchmarks in simulated and production environments.
  • Create datasets and scenarios to evaluate model performance under constraints.
  • Identify bottlenecks and implement solutions including batching and memory management.
  • Develop GPU kernels for mobile hardware and shaders in MSL.
  • Apply pruning, quantization, and advanced inference optimization techniques.
  • Design distributed inference using tensor/pipeline/expert parallelism.
  • Collaborate with cross-functional teams to productionize optimized frameworks.
  • Define evaluation methodologies and document experimental results.
  • Monitor production performance to drive further improvements.

Kenntnisse

MSL kernel optimization
GPU kernel development
Model serving architectures
Inference latency optimization
Benchmarking and evaluation

Ausbildung

PhD in NLP/ML or related with AI research track

Tools

Metal Shading Language (MSL)
Tensor parallelism

Jobbeschreibung

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in Germany.

Type: Onsite

You will work at the intersection of AI research, systems engineering, and high-performance model inference.
Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments.
You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices.
The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels.
You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers.
Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements.
You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems.

Accountabilities
  • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
  • Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
  • Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
  • Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
  • Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
  • Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
  • Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
  • Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads.
  • Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications.
  • Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies.
  • Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability.
Requirements
  • Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences.
  • Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch.
  • Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices.
  • Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications.
  • Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment.
  • Strong experience writing GPU kernels for mobile devices such as smartphones.
  • Practical experience developing and deploying end-to-end inference pipelines, from model optimization through production integration on constrained hardware.
  • Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges.
  • Experience designing robust evaluation and benchmarking frameworks for inference systems.
  • Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters.
  • Deep understanding of the mathematical foundations and architecture of diffusion models and Vision Transformers.
  • Familiarity with modern inference optimization techniques including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE.
  • Strong analytical and problem-solving abilities, with an ability to investigate complex system bottlenecks and turn research findings into practical engineering solutions.
  • Excellent English communication skills and the ability to collaborate effectively with distributed, cross-functional technical teams.
Benefits
  • Opportunity to work on advanced AI systems spanning model serving, inference optimization, mobile computing, edge deployment, and large-scale distributed inference.
  • Remote-first working environment with an international team.
  • Exposure to cutting-edge AI research and practical systems engineering challenges.
  • Opportunity to contribute to performance-critical infrastructure where improvements can have a measurable impact on real-world AI applications.
  • Collaborative environment combining research-driven experimentation with hands-on engineering.
  • Opportunity to work with advanced model architectures including diffusion models, Vision Transformers, and multimodal systems.
  • Access to challenging technical problems involving GPU kernels, inference engines, memory optimization, and distributed computing.

This overview was written from the original listing to help you understand the role. Always confirm the specifics on the employer's application page.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

United States Digital Space LLC • Deutschland

Remote
EUR 120.000 - 190.000
Remote-first
International team
Cutting-edge AI research
+2
AI Architect(m/f/d)
AI Architect(m/f/d)

NC GROUP GmbH • Berlin

Vor Ort
EUR 80.000 - 120.000
28 vacation days
Health & Sports Subsidy
Public Transportation Subsidy
+2
AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide
AI Research Engineer (Model Compression & Quantization) - 100% Remote Worldwide

Tether Operations Limited • Deutschland

Remote
EUR 90.000 - 130.000
Member of Technical Staff - Inference & Hardware Optimization
Member of Technical Staff - Inference & Hardware Optimization

Albs Labs GmbH • Freiburg im Breisgau

Hybrid
EUR 90.000 - 130.000
Senior AI Platform & Research Infrastructure Engineer (m/f/d)
Senior AI Platform & Research Infrastructure Engineer (m/f/d)

United States Digital Space LLC • Garching bei München

Hybrid
EUR 110.000 - 170.000
60% remote work option
Free parking
Pension plan (VBL)
+1
Senior AI Expert (m/f/d) for AI Computing Services
Senior AI Expert (m/f/d) for AI Computing Services

Leibniz-Rechenzentrum • Garching bei München

Vor Ort
EUR 58.500 - 71.500
Pension plan
State-of-the-art work equipment
Free parking
Applied AI Scientist(m/f/d)
Applied AI Scientist(m/f/d)

Nc Group • Berlin

Vor Ort
EUR 90.000 - 140.000
Full-time
Flexible working hours
28 vacation days
+4
Radar Signal Processing & AI Engineer - Ground-Up Systems
Radar Signal Processing & AI Engineer - Ground-Up Systems

THRYVE • Berlin

Vor Ort
EUR 120.000 - 190.000
Competitive compensation
Stock options
Flexible work environment
+1
Forward Deployed Engineer – AI Inference (Intern)
Forward Deployed Engineer – AI Inference (Intern)

Lyceum • Berlin

Vor Ort
EUR 8.900 - 13.000
Senior Machine Learning Research Developer
Senior Machine Learning Research Developer

European Tech Recruit • Berlin

Hybrid
EUR 90.000 - 130.000