Remote AI Inference Engineer - Kernel & Performance

Lever, Inc.

Saudi Arabia

On-site

SAR 420,000 - 600,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Remote-first environment
International team
Cutting-edge AI research exposure

Job summary

Lever, Inc. partners with a specialized company to search for an AI Research Engineer (Kernel & Inference Optimization) in Saudi Arabia.

The role sits at the intersection of AI research, systems engineering, and high-performance model inference, focusing on optimizing model-serving architectures for diverse hardware, including edge devices. You will tackle latency, throughput, and memory challenges, develop GPU kernels in Metal Shading Language, and apply pruning, quantization, and Flash

Qualifications

  • Degree in Computer Science or related field; a PhD in NLP, ML or related discipline with publications is highly relevant.
  • Proven expertise in Metal Shading Language (MSL) and writing custom compute shaders from scratch.
  • Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices.
  • Track record of delivering measurable improvements in latency, throughput, and memory footprint for domain-specific applications.
  • Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment.
  • Strong experience writing GPU kernels for mobile devices.
  • Practical experience developing and deploying end-to-end inference pipelines on constrained hardware.
  • Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges.
  • Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters.
  • Deep understanding of diffusion models and Vision Transformers.
  • Familiarity with pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE.
  • Strong analytical and problem-solving abilities and excellent English communication skills.

Responsibilities

  • Design and deploy advanced model-serving architectures for high throughput and low latency.
  • Develop inference pipelines across mobile and edge platforms.
  • Set performance targets for latency, token generation speed, throughput, memory footprint, and reliability.
  • Build controlled inference benchmarks and track latency, throughput, memory, and errors.
  • Create datasets and simulation scenarios for evaluating model performance under real-world constraints.
  • Identify bottlenecks across inference pipelines and implement batching, memory management, and other optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware (MSL).
  • Apply pruning, quantization, Flash Attention, KV caching, and other optimization techniques.
  • Design distributed inference systems using tensor/pipeline/expert parallelism for large GPU workloads.
  • Collaborate with cross-functional teams to integrate inference frameworks into production and edge-device apps.
  • Define evaluation methodologies and document experimental results against benchmarks.
  • Monitor production performance and pursue scalability improvements.

Skills

GPU kernel optimization
Low-level optimization
Inference optimization
Distributed inference
Benchmarking
Empirical research
Cross-functional collaboration

Education

PhD in NLP or ML
Degree in Computer Science or related field

Tools

Metal Shading Language (MSL)
Flash Attention
KV Cache optimization

Job description

Lever, Inc. partners with a specialized company to search for an AI Research Engineer (Kernel & Inference Optimization) in Saudi Arabia.

The role sits at the intersection of AI research, systems engineering, and high-performance model inference, focusing on optimizing model-serving architectures for diverse hardware, including edge devices. You will tackle latency, throughput, and memory challenges, develop GPU kernels in Metal Shading Language, and apply pruning, quantization, and Flash

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Enterprise AI Platform Engineer | LLM & GPU Inference
Enterprise AI Platform Engineer | LLM & GPU Inference

Saudi Aramco (ASC) • Saudi Arabia

On-site
SAR 299,962 - 449,943
AI Engineer - LLMs & Arabic NLP, Production Infra
AI Engineer - LLMs & Arabic NLP, Production Infra

InnovationTeam • Riyadh

On-site
SAR 200,000 - 420,000
Senior ML Platform Engineer, Inference & MLOps
Senior ML Platform Engineer, Inference & MLOps

Mollkom • Riyadh

Hybrid
SAR 150,000 - 210,000
Ai/ml/llm Systems Engineer - Enterprise Ai Platform Engineer
Ai/ml/llm Systems Engineer - Enterprise Ai Platform Engineer

Saudi Aramco (ASC) • Saudi Arabia

On-site
SAR 299,962 - 449,943
AI-Native Product Engineer: End-to-End, Remote
AI-Native Product Engineer: End-to-End, Remote

Lever, Inc. • Saudi Arabia

Remote
SAR 180,000 - 360,000
Fully remote work
Vacation 28 days per year
Wellness days 7 days per year
+3
Senior ML Platform Engineer - Inference & MLOps (Hybrid)
Senior ML Platform Engineer - Inference & MLOps (Hybrid)

Mollkom • Riyadh

Hybrid
SAR 150,000 - 210,000
Remote AI Process Deployment Engineer
Remote AI Process Deployment Engineer

Lever, Inc. • Saudi Arabia

On-site
SAR 180,000 - 300,000
Fully remote work
Home office budget
Flexible time off
+6
Senior Quality Engineer - AI-Driven Testing & Automation
Senior Quality Engineer - AI-Driven Testing & Automation

Lever, Inc. • Saudi Arabia

On-site
SAR 539,000 - 870,000
Fully remote
Home office budget
Flexible time off
+3
AI Delivery Engineer: Deploy & Scale ML Models
AI Delivery Engineer: Deploy & Scale ML Models

Elm Co • Riyadh

On-site
SAR 180,000 - 300,000
Associate Data Scientist (Saudi Arabia)
Associate Data Scientist (Saudi Arabia)

Eram Talent • Dhahran Compound

On-site
SAR 180,000 - 300,000