AI Inference Engineer - Kernel & Edge Optimization

Jobgether SRL

United States

Remote

USD 82,000 - 177,000

Full time

42 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Remote-first team
Exposure to cutting-edge AI research
Collaborative engineering environment

Job summary

Jobgether SRL is seeking an AI Research Engineer (Kernel & Inference Optimization) to advance model-serving architectures for diverse hardware and edge devices. You will blend hands-on research with low-level engineering to push latency, throughput, and memory efficiency while targeting mobile and edge platforms.

Collaborating with cross-functional, remote teams, you will benchmark, test, and translate research into scalable production solutions, including diffusion models, Vision Transformers,

Qualifications

  • PhD in NLP/ML or related field with AI research track record.
  • Proven expertise with MSL and custom compute shaders.
  • Experience with low-level kernel and inference optimization on mobile/resource-constrained devices.
  • Track record of improving inference latency, throughput, and memory footprint.
  • Strong understanding of model-serving architectures and optimization techniques.
  • Experience writing GPU kernels for mobile devices.
  • Proven ability to develop end-to-end inference pipelines for production.
  • Strong empirical research and systematic experimentation capabilities.
  • Familiarity with distributed inference techniques for large-scale GPUs.

Responsibilities

  • Design and deploy advanced model-serving architectures for high throughput and low latency.
  • Develop inference pipelines across diverse environments including edge devices.
  • Set clear performance targets for latency, throughput, memory, and reliability.
  • Create controlled benchmarks to track latency, throughput, memory use, and errors.
  • Develop datasets and simulations for evaluating performance.
  • Identify bottlenecks and implement system-level optimizations.
  • Build custom GPU kernels and compute shaders for mobile hardware (MSL).
  • Apply pruning, quantization, Flash Attention, KV caching, and speculative decoding.
  • Design distributed inference systems using tensor/pipeline/expert parallelism.
  • Collaborate to integrate optimized inference into production and edge apps.
  • Document experiments and refine optimization strategies based on results.
  • Monitor production performance and identify scalability improvements.

Skills

Low-level kernel optimization
Inference optimization
Empirical research
English communication
System design
GPU performance tuning

Education

PhD in NLP/ML or related field
Master's degree in CS

Tools

Metal Shading Language (MSL)
GPU kernels
Inference frameworks
Quantization

Job description

Jobgether SRL is seeking an AI Research Engineer (Kernel & Inference Optimization) to advance model-serving architectures for diverse hardware and edge devices. You will blend hands-on research with low-level engineering to push latency, throughput, and memory efficiency while targeting mobile and edge platforms.

Collaborating with cross-functional, remote teams, you will benchmark, test, and translate research into scalable production solutions, including diffusion models, Vision Transformers,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Engineer: Kernel & Edge Optimization
AI Inference Engineer: Kernel & Edge Optimization

Lever, Inc. • Town of Italy (NY)

On-site
EUR 90,000 - 130,000
Remote-first environment
International team
Exposure to cutting-edge AI research
+1
AI Inference Kernel Optimization Engineer
AI Inference Kernel Optimization Engineer

Lever, Inc. • Town of Belgium (WI)

Remote
USD 102,000 - 147,000
Remote-first
International team
Cutting-edge AI research
+1
Edge AI Inference Architect: Kernel & Performance
Edge AI Inference Architect: Kernel & Performance

Lever, Inc. • Spain (TX)

Remote
USD 136,000 - 204,000
Senior AI Performance Engineer - Edge Inference
Senior AI Performance Engineer - Edge Inference

Arm Limited • San Jose (CA)

On-site
USD 263,000 - 355,000
Senior AI Performance Engineer — Edge Inference Expert
Senior AI Performance Engineer — Edge Inference Expert

Arm • San Jose (CA)

Hybrid
USD 263,000 - 355,000
AI Kernel Engineer — Optimize LLM Inference & Kernels
AI Kernel Engineer — Optimize LLM Inference & Kernels

quadric • Burlingame (CA)

On-site
USD 170,000 - 230,000
Medical, dental, and vision insurance
Equity
Discretionary annual bonus
+3
Edge AI Kernel Engineer - High-Performance Inference
Edge AI Kernel Engineer - High-Performance Inference

QUADRIC PTY LTD • Burlingame (CA)

On-site
USD 170,000 - 230,000
Medical, dental, and vision insurance
Company-paid life insurance
Voluntary life insurance
+11
Senior AI Kernel Engineer — Edge Inference & Optimization
Senior AI Kernel Engineer — Edge Inference & Optimization

quadric.io, Inc • Burlingame (CA)

On-site
USD 120,000 - 150,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Life Insurance (Basic, Voluntary & AD&D)
+7
Senior AI Kernel Engineer - Edge ML Optimization
Senior AI Kernel Engineer - Edge ML Optimization

Quadric Inc. • Burlingame (CA)

Hybrid
USD 110,000 - 270,000
Competitive salary
Equity
Health, dental, and vision
+7
AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Town of Italy (NY)

On-site
EUR 90,000 - 130,000
Remote-first environment
International team
Exposure to cutting-edge AI research
+1