AI Inference Kernel Optimization Engineer

Lever, Inc.

Town of Belgium (WI)

Remote

USD 102,000 - 147,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Remote-first
International team
Cutting-edge AI research
Performance-critical infrastructure

Job summary

Lever, Inc. is seeking an AI Research Engineer focused on Kernel & Inference Optimization, based in Belgium with a remote-first, international team. The role spans AI research, low-level engineering, and high-performance model serving across devices from mobile to edge.

You will tackle latency, throughput, and memory challenges, develop GPU kernels (MSL), and advance inference techniques like pruning and quantization. Collaboration with research and engineering teams is essential.

Qualifications

  • PhD in NLP/ML or related field with AI research track record preferred.
  • Proven expertise in writing compute shaders and MSL.
  • Experience with low-level kernel optimization on mobile or constrained devices.
  • Track record of improvements in latency, throughput, and memory footprint for AI applications.
  • Deep understanding of model-serving architectures and optimization techniques.
  • Experience writing GPU kernels for mobile devices.
  • End-to-end inference pipelines from optimization to production on constrained hardware.
  • Ability to apply empirical research and systematic experimentation to optimize latency and memory.

Responsibilities

  • Design and deploy model-serving architectures with high throughput and low latency.
  • Develop inference pipelines across mobile and edge platforms.
  • Set performance targets for latency, throughput, memory, and reliability.
  • Build and run inference benchmarks in simulated and production environments.
  • Create datasets and scenarios for evaluating performance under resource constraints.
  • Identify bottlenecks and implement batching, memory management, and networking optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware (MSL).
  • Apply pruning, quantization, Flash Attention, KV caching, and speculative decoding to optimize inference.
  • Design distributed inference using tensor/pipeline/expert parallelism for large GPUs.
  • Collaborate with cross-functional teams to integrate optimized inference frameworks into production.
  • Document evaluation methodologies and compare to benchmarks; iterate on optimization strategies.
  • Monitor production performance and seek continual improvements in scalability and efficiency.

Skills

MSL expertise
Low-level kernel optimization
Inference optimization
GPU kernels
Distributed inference
Diffusion models knowledge
Vision Transformers knowledge
English communication

Education

PhD in NLP/ML (relevant)
Bachelor's degree in CS or related field

Tools

Metal Shading Language (MSL)
Kernel development tools
Edge devices tooling

Job description

Lever, Inc. is seeking an AI Research Engineer focused on Kernel & Inference Optimization, based in Belgium with a remote-first, international team. The role spans AI research, low-level engineering, and high-performance model serving across devices from mobile to edge.

You will tackle latency, throughput, and memory challenges, develop GPU kernels (MSL), and advance inference techniques like pruning and quantization. Collaboration with research and engineering teams is essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Engineer: Kernel & Edge Optimization
AI Inference Engineer: Kernel & Edge Optimization

Lever, Inc. • Town of Italy (NY)

On-site
EUR 90,000 - 130,000
Remote-first environment
International team
Exposure to cutting-edge AI research
+1
Edge AI Inference Architect: Kernel & Performance
Edge AI Inference Architect: Kernel & Performance

Lever, Inc. • Spain (TX)

Remote
USD 136,000 - 204,000
AI Inference Engineer - Kernel & Edge Optimization
AI Inference Engineer - Kernel & Edge Optimization

Jobgether SRL • United States

Remote
USD 82,000 - 177,000
Remote-first team
Exposure to cutting-edge AI research
Collaborative engineering environment
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Town of Belgium (WI)

Remote
USD 102,000 - 147,000
Remote-first
International team
Cutting-edge AI research
+1
Senior AI Kernel Engineer — Remote/Hybrid Inference
Senior AI Kernel Engineer — Remote/Hybrid Inference

Modular • United States

Hybrid
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
AI Kernel Engineer — Optimize LLM Inference & Kernels
AI Kernel Engineer — Optimize LLM Inference & Kernels

quadric • Burlingame (CA)

On-site
USD 170,000 - 230,000
Medical, dental, and vision insurance
Equity
Discretionary annual bonus
+3
Senior ML Research Engineer - Inference & GPU Kernels
Senior ML Research Engineer - Inference & GPU Kernels

AfterQuery • Tempe (AZ)

Remote
USD 138,000 - 234,000
Frontier AI work
Fully remote
Competitive pay
AI Research Engineer: Inference & GPU Kernels (Remote)
AI Research Engineer: Inference & GPU Kernels (Remote)

AfterQuery • Bridgewater (MA)

Remote
USD 157,000 - 215,000
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Spain (TX)

Remote
USD 135,000 - 203,000
Remote-first
International team
Cutting-edge AI
+1
Senior AI Inference & GPU Kernel Research Engineer (Remote)
Senior AI Inference & GPU Kernel Research Engineer (Remote)

AfterQuery • Lansing (MI)

Remote
USD 158,000 - 214,000