Edge AI Inference Architect: Kernel & Performance

Lever, Inc.

Spain (TX)

Remote

USD 136,000 - 204,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Lever, Inc. is seeking an AI Research Engineer (Kernel & Inference Optimization) in Spain to push the boundaries of efficient AI systems.

You will work at the intersection of AI research, systems engineering, and high-performance model inference, delivering practical improvements across diverse hardware environments. You will develop and optimize model-serving architectures, write custom GPU kernels for mobile hardware (MSL), and apply pruning, quantization, and Flash Attention to reduce latency

Qualifications

  • PhD in NLP, ML or related field with strong AI research track record.
  • Experience in low-level kernel optimization for mobile or edge devices.
  • Proven ability to translate research into optimized inference for production.

Responsibilities

  • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and memory efficiency.
  • Develop inference pipelines across mobile and edge platforms.
  • Benchmark and document performance improvements against benchmarks.
  • Create datasets and scenarios for evaluating model performance under constraints.
  • Identify bottlenecks and implement system-level optimizations.

Skills

Model serving
Latency optimization
GPU kernels
Metal Shading Language
Empirical benchmarking
Distributed inference
Research-to-production

Education

PhD in NLP/ML
CS degree

Tools

MSL

Job description

Lever, Inc. is seeking an AI Research Engineer (Kernel & Inference Optimization) in Spain to push the boundaries of efficient AI systems.

You will work at the intersection of AI research, systems engineering, and high-performance model inference, delivering practical improvements across diverse hardware environments. You will develop and optimize model-serving architectures, write custom GPU kernels for mobile hardware (MSL), and apply pruning, quantization, and Flash Attention to reduce latency

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Engineer: Kernel & Edge Optimization
AI Inference Engineer: Kernel & Edge Optimization

Lever, Inc. • Town of Italy (NY)

On-site
EUR 90,000 - 130,000
Remote-first environment
International team
Exposure to cutting-edge AI research
+1
AI Inference Kernel Optimization Engineer
AI Inference Kernel Optimization Engineer

Lever, Inc. • Town of Belgium (WI)

Remote
USD 102,000 - 147,000
Remote-first
International team
Cutting-edge AI research
+1
AI Inference Engineer - Kernel & Edge Optimization
AI Inference Engineer - Kernel & Edge Optimization

Jobgether SRL • United States

Remote
USD 82,000 - 177,000
Remote-first team
Exposure to cutting-edge AI research
Collaborative engineering environment
Kernel Engineer: GPU Performance & Inference
Kernel Engineer: GPU Performance & Inference

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Edge AI Kernel Engineer - High-Performance Inference
Edge AI Kernel Engineer - High-Performance Inference

QUADRIC PTY LTD • Burlingame (CA)

On-site
USD 170,000 - 230,000
Medical, dental, and vision insurance
Company-paid life insurance
Voluntary life insurance
+11
Senior Inference Engineer: AI-Driven GPU Kernel Optimization
Senior Inference Engineer: AI-Driven GPU Kernel Optimization

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Spain (TX)

Remote
USD 135,000 - 203,000
Remote-first
International team
Cutting-edge AI
+1
Senior AI Kernel Engineer — Remote/Hybrid Inference
Senior AI Kernel Engineer — Remote/Hybrid Inference

Modular • United States

Hybrid
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior AI Performance Engineer - Edge Inference
Senior AI Performance Engineer - Edge Inference

Arm Limited • San Jose (CA)

On-site
USD 263,000 - 355,000