AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc.

Town of Belgium (WI)

Remote

USD 102,000 - 147,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Remote-first
International team
Cutting-edge AI research
Performance-critical infrastructure

Job summary

Lever, Inc. is seeking an AI Research Engineer focused on Kernel & Inference Optimization, based in Belgium with a remote-first, international team. The role spans AI research, low-level engineering, and high-performance model serving across devices from mobile to edge.

You will tackle latency, throughput, and memory challenges, develop GPU kernels (MSL), and advance inference techniques like pruning and quantization. Collaboration with research and engineering teams is essential.

Qualifications

  • PhD in NLP/ML or related field with AI research track record preferred.
  • Proven expertise in writing compute shaders and MSL.
  • Experience with low-level kernel optimization on mobile or constrained devices.
  • Track record of improvements in latency, throughput, and memory footprint for AI applications.
  • Deep understanding of model-serving architectures and optimization techniques.
  • Experience writing GPU kernels for mobile devices.
  • End-to-end inference pipelines from optimization to production on constrained hardware.
  • Ability to apply empirical research and systematic experimentation to optimize latency and memory.

Responsibilities

  • Design and deploy model-serving architectures with high throughput and low latency.
  • Develop inference pipelines across mobile and edge platforms.
  • Set performance targets for latency, throughput, memory, and reliability.
  • Build and run inference benchmarks in simulated and production environments.
  • Create datasets and scenarios for evaluating performance under resource constraints.
  • Identify bottlenecks and implement batching, memory management, and networking optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware (MSL).
  • Apply pruning, quantization, Flash Attention, KV caching, and speculative decoding to optimize inference.
  • Design distributed inference using tensor/pipeline/expert parallelism for large GPUs.
  • Collaborate with cross-functional teams to integrate optimized inference frameworks into production.
  • Document evaluation methodologies and compare to benchmarks; iterate on optimization strategies.
  • Monitor production performance and seek continual improvements in scalability and efficiency.

Skills

MSL expertise
Low-level kernel optimization
Inference optimization
GPU kernels
Distributed inference
Diffusion models knowledge
Vision Transformers knowledge
English communication

Education

PhD in NLP/ML (relevant)
Bachelor's degree in CS or related field

Tools

Metal Shading Language (MSL)
Kernel development tools
Edge devices tooling

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in Belgium.

You will work at the intersection of AI research, systems engineering, and high-performance model inference.
Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments.
You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices.
The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels.
You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers.
Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements.
You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems.

Accountabilities
  • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
  • Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
  • Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
  • Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
  • Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
  • Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
  • Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
  • Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads.
  • Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications.
  • Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies.
  • Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability.
Requirements:
  • Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences.
  • Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch.
  • Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices.
  • Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications.
  • Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment.
  • Strong experience writing GPU kernels for mobile devices such as smartphones.
  • Practical experience developing and deploying end-to-end inference pipelines, from model optimization through production integration on constrained hardware.
  • Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges.
  • Experience designing robust evaluation and benchmarking frameworks for inference systems.
  • Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters.
  • Deep understanding of the mathematical foundations and architecture of diffusion models and Vision Transformers.
  • Familiarity with modern inference optimization techniques including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE.
  • Strong analytical and problem-solving abilities, with an ability to investigate complex system bottlenecks and turn research findings into practical engineering solutions.
  • Excellent English communication skills and the ability to collaborate effectively with distributed, cross-functional technical teams.
Benefits:
  • Opportunity to work on advanced AI systems spanning model serving, inference optimization, mobile computing, edge deployment, and large-scale distributed inference.
  • Remote-first working environment with an international team.
  • Exposure to cutting-edge AI research and practical systems engineering challenges.
  • Opportunity to contribute to performance-critical infrastructure where improvements can have a measurable impact on real-world AI applications.
  • Collaborative environment combining research-driven experimentation with hands-on engineering.
  • Opportunity to work with advanced model architectures including diffusion models, Vision Transformers, and multimodal systems.
  • Access to challenging technical problems involving GPU kernels, inference engines, memory optimization, and distributed computing.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Town of Italy (NY)

On-site
EUR 90,000 - 130,000
Remote-first environment
International team
Exposure to cutting-edge AI research
+1
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Spain (TX)

Remote
USD 135,000 - 203,000
Remote-first
International team
Cutting-edge AI
+1
AI Research Engineer, Inference
AI Research Engineer, Inference

OP Recruiting • Chicago (IL)

On-site
USD 150,000 - 210,000
Health benefits
Research Engineer - Inference
Research Engineer - Inference

ElevenLabs • Maine

On-site
USD 140,000 - 190,000
Annual discretionary stipend
Annual company offsite
Co-working stipend
AI Inference Kernel Optimization Engineer
AI Inference Kernel Optimization Engineer

Lever, Inc. • Town of Belgium (WI)

Remote
USD 102,000 - 147,000
Remote-first
International team
Cutting-edge AI research
+1
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Senior AI Engineer (f/m/x)
Senior AI Engineer (f/m/x)

Lever, Inc. • Germany (OH)

Remote
USD 136,000 - 181,000
Fully remote-first environment
Vienna office option
Flexible hours
+5
Applied Researcher – AI Expert
Applied Researcher – AI Expert

Designworks Talent • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
AI Inference Engineer (f/m/d)
AI Inference Engineer (f/m/d)

ITRex Group • Town of Poland (NY)

On-site
USD 100,000 - 130,000
Remote flexibility
Competitive salary and medical benefits
Career progression opportunities
AI Infra Staff Researcher
AI Infra Staff Researcher

Lenovo • North Carolina

On-site
USD 140,000 - 190,000