AI Inference Engineer: Kernel & Edge Optimization

Lever, Inc.

Town of Italy (NY)

On-site

EUR 90,000 - 130,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Remote-first environment
International team
Exposure to cutting-edge AI research
Performance-critical infrastructure

Job summary

Lever, Inc. is seeking an AI Research Engineer (Kernel & Inference Optimization) based in Italy.

You will work at the intersection of AI research, systems engineering, and high-performance model inference, developing and optimizing model-serving architectures for diverse hardware environments. You will tackle latency, throughput, memory efficiency and scalability challenges across mobile and edge platforms, including GPU kernel work and advanced optimization techniques.

Qualifications

  • PhD in NLP/ML or related field or strong AI research track record.
  • Proven expertise in Metal Shading Language (MSL) and writing custom compute shaders.
  • Experience with low-level kernel optimization and inference optimization on mobile or constrained devices.
  • Track record of improving inference latency, throughput and memory footprint.
  • Strong understanding of model-serving architectures and optimization techniques.
  • Experience developing GPU kernels for mobile devices.
  • Ability to run empirical research and systematic experiments to solve bottlenecks.

Responsibilities

  • Design and deploy model-serving architectures optimized for throughput and latency.
  • Develop inference pipelines for mobile and edge platforms.
  • Set performance targets for latency, speed, throughput, memory and reliability.
  • Create benchmarks and collect metrics on latency, throughput and errors.
  • Build datasets and simulations for evaluating performance.
  • Identify bottlenecks and implement batching, memory management and networking optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware (MSL).
  • Apply pruning, quantization, Flash Attention, KV caching and speculative decoding.
  • Design distributed inference with tensor/pipeline/expert parallelism for large GPUs.
  • Collaborate to integrate optimized frameworks into production and edge devices.
  • Define evaluation methodologies and report experiments against benchmarks.
  • Monitor production performance for scalability and reliability improvements.

Skills

MSL programming
GPU kernel optimization
Inference optimization
Tensor/parallels
Diffusion models
Vision Transformers
Edge/mobile deployment
Benchmarking

Education

PhD in NLP/ML or related
Degree in Computer Science

Tools

TensorRT/ONNX Runtime
MLIR
CoreML/Metal tooling

Job description

Lever, Inc. is seeking an AI Research Engineer (Kernel & Inference Optimization) based in Italy.

You will work at the intersection of AI research, systems engineering, and high-performance model inference, developing and optimizing model-serving architectures for diverse hardware environments. You will tackle latency, throughput, memory efficiency and scalability challenges across mobile and edge platforms, including GPU kernel work and advanced optimization techniques.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Engineer - Kernel & Edge Optimization
AI Inference Engineer - Kernel & Edge Optimization

Jobgether SRL • United States

Remote
USD 82,000 - 177,000
Remote-first team
Exposure to cutting-edge AI research
Collaborative engineering environment
Edge AI Inference Architect: Kernel & Performance
Edge AI Inference Architect: Kernel & Performance

Lever, Inc. • Spain (TX)

Remote
USD 136,000 - 204,000
AI Inference Kernel Optimization Engineer
AI Inference Kernel Optimization Engineer

Lever, Inc. • Town of Belgium (WI)

Remote
USD 102,000 - 147,000
Remote-first
International team
Cutting-edge AI research
+1
Edge AI Kernel Engineer - High-Performance Inference
Edge AI Kernel Engineer - High-Performance Inference

QUADRIC PTY LTD • Burlingame (CA)

On-site
USD 170,000 - 230,000
Medical, dental, and vision insurance
Company-paid life insurance
Voluntary life insurance
+11
AI Kernel Engineer — Optimize LLM Inference & Kernels
AI Kernel Engineer — Optimize LLM Inference & Kernels

quadric • Burlingame (CA)

On-site
USD 170,000 - 230,000
Medical, dental, and vision insurance
Equity
Discretionary annual bonus
+3
Senior AI Performance Engineer - Edge Inference
Senior AI Performance Engineer - Edge Inference

Arm Limited • San Jose (CA)

On-site
USD 263,000 - 355,000
Senior AI Kernel Engineer - Edge ML Optimization
Senior AI Kernel Engineer - Edge ML Optimization

Quadric Inc. • Burlingame (CA)

Hybrid
USD 110,000 - 270,000
Competitive salary
Equity
Health, dental, and vision
+7
Senior AI Performance Engineer — Edge Inference Expert
Senior AI Performance Engineer — Edge Inference Expert

Arm • San Jose (CA)

Hybrid
USD 263,000 - 355,000
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Town of Italy (NY)

On-site
EUR 90,000 - 130,000
Remote-first environment
International team
Exposure to cutting-edge AI research
+1
Edge AI Kernel Engineer - Real-Time Inference
Edge AI Kernel Engineer - Real-Time Inference

Quadric • Burlingame (CA)

Hybrid
CAD 155,000 - 382,000
Competitive salary
Meaningful equity
Medical/Dental/Vision
+2