AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc.

Schweiz

Remote

EUR 137.000 - 222.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Verschicke keinen 08/15-Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Remote-first team
International team
Cutting-edge AI research
Performance-critical infra
Collaboration with researchers
Edge deployment challenges

Zusammenfassung

Lever, Inc. is seeking an AI Research Engineer (Kernel & Inference Optimization) based in Switzerland. You will push the boundaries of efficient AI systems across hardware environments, balancing latency, throughput, and memory in model-serving architectures.

You will research and implement advanced inference strategies, develop custom GPU kernels, and collaborate across teams to translate research into production-ready infrastructure for edge devices.

Qualifikationen

  • PhD in NLP/ML with strong AI research track record and publications.
  • Proven expertise in MSL and writing custom compute shaders.
  • Experience with low-level kernel optimization on mobile/resource-constrained devices.
  • Track record delivering measurable improvements in latency, throughput and memory.
  • Deep understanding of model-serving architectures and optimization techniques.
  • Experience writing GPU kernels for mobile devices.
  • End-to-end inference pipelines from model optimization to production on constrained hardware.
  • Ability to apply empirical research and systematic experimentation to solve bottlenecks.
  • Experience benchmarking inference systems across scales and environments.
  • Knowledge of distributed inference: tensor/pipeline/expert parallelism for large GPUs.
  • Understanding diffusion models and Vision Transformers architecture.
  • Familiarity with pruning, quantization, Flash Attention, KV Cache, speculative decoding.

Aufgaben

  • Design and deploy advanced model-serving architectures for high throughput and low latency.
  • Develop inference pipelines for mobile and edge platforms with constrained resources.
  • Set performance targets for latency, throughput, memory and reliability.
  • Build controlled benchmarks and collect latency, throughput and memory metrics.
  • Create datasets and simulations to evaluate model performance in real-world conditions.
  • Identify bottlenecks and implement solutions in batching, networking and memory management.
  • Develop custom GPU kernels and compute shaders for mobile hardware (MSL).
  • Apply inference optimization techniques including pruning, quantization and Flash Attention.
  • Design distributed inference with tensor/pipeline/expert parallelism for large GPU workloads.
  • Collaborate with cross-functional teams to integrate optimized inference into production and edge apps.
  • Document results and refine optimization strategies based on empirical evidence.
  • Monitor production performance and seek improvements in scalability and efficiency.

Kenntnisse

Latency optimization
Benchmarking
Empirical research
English communication
Tensor parallelism
Distributed inference

Ausbildung

PhD in NLP or ML
CS degree

Tools

Metal Shading Language (MSL)

Jobbeschreibung

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in Switzerland.

You will work at the intersection of AI research, systems engineering, and high-performance model inference.
Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments.
You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices.
The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels.
You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers.
Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements.
You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems.

Accountabilities
  • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
  • Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
  • Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
  • Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
  • Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
  • Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
  • Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
  • Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads.
  • Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications.
  • Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies.
  • Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability.
Requirements:
  • Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences.
  • Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch.
  • Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices.
  • Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications.
  • Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment.
  • Strong experience writing GPU kernels for mobile devices such as smartphones.
  • Practical experience developing and deploying end-to-end inference pipelines, from model optimization through production integration on constrained hardware.
  • Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges.
  • Experience designing robust evaluation and benchmarking frameworks for inference systems.
  • Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters.
  • Deep understanding of the mathematical foundations and architecture of diffusion models and Vision Transformers.
  • Familiarity with modern inference optimization techniques including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE.
  • Strong analytical and problem-solving abilities, with an ability to investigate complex system bottlenecks and turn research findings into practical engineering solutions.
  • Excellent English communication skills and the ability to collaborate effectively with distributed, cross-functional technical teams.
Benefits:
  • Opportunity to work on advanced AI systems spanning model serving, inference optimization, mobile computing, edge deployment, and large-scale distributed inference.
  • Remote-first working environment with an international team.
  • Exposure to cutting-edge AI research and practical systems engineering challenges.
  • Opportunity to contribute to performance-critical infrastructure where improvements can have a measurable impact on real-world AI applications.
  • Collaborative environment combining research-driven experimentation with hands-on engineering.
  • Opportunity to work with advanced model architectures including diffusion models, Vision Transformers, and multimodal systems.
  • Access to challenging technical problems involving GPU kernels, inference engines, memory optimization, and distributed computing.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior AI Inference Engineer — Kernel & Edge Optimization
Senior AI Inference Engineer — Kernel & Edge Optimization

Lever, Inc. • Schweiz

Remote
EUR 137.000 - 222.000
Remote-first team
International team
Cutting-edge AI research
+3
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Schweiz

Vor Ort
CHF 100.000 - 140.000
AI Research Scientist
AI Research Scientist

Embodied AI • Lausanne

Vor Ort
CHF 120.000 - 170.000
AI Engineer - Switzerland (Remote)
AI Engineer - Switzerland (Remote)

Yeah! Global • Zürich

Remote
CHF 100.000 - 130.000
AI Engineer
AI Engineer

Artificialy • Zürich

Vor Ort
CHF 110.000 - 150.000
Full-time contract
Mentorship and continuous learning
Office in Zurich or Lugano
Senior AI/ML Engineer, CH
Senior AI/ML Engineer, CH

vector8 • Zürich

Vor Ort
CHF 90.000 - 120.000
Competitive salary package
25 days of vacation
Development budget
+1
Research Scientist, AI RL and LLMs
Research Scientist, AI RL and LLMs

NVIDIA Gruppe • Zürich

Vor Ort
CHF 180.000 - 260.000
Comprehensive benefits
Competitive salary
AI Developer Technology Engineer
AI Developer Technology Engineer

NVIDIA • Zürich

Vor Ort
CHF 140.000 - 210.000
AI Science Writer, Nebius Academy (Contract)
AI Science Writer, Nebius Academy (Contract)

Jobgether SRL • Schweiz

Remote
CHF 78.000 - 106.000
Fixed monthly retainer
European remote work options
Career growth opportunities
+1
Applied ML Engineer
Applied ML Engineer

Jobgether SRL • Schweiz

Vor Ort
CHF 31.000 - 52.000