Senior AI Inference Engineer — Kernel & Edge Optimization

Lever, Inc.

Schweiz

Remote

EUR 137.000 - 222.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Verschicke keinen 08/15-Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Remote-first team
International team
Cutting-edge AI research
Performance-critical infra
Collaboration with researchers
Edge deployment challenges

Zusammenfassung

Lever, Inc. is seeking an AI Research Engineer (Kernel & Inference Optimization) based in Switzerland. You will push the boundaries of efficient AI systems across hardware environments, balancing latency, throughput, and memory in model-serving architectures.

You will research and implement advanced inference strategies, develop custom GPU kernels, and collaborate across teams to translate research into production-ready infrastructure for edge devices.

Qualifikationen

  • PhD in NLP/ML with strong AI research track record and publications.
  • Proven expertise in MSL and writing custom compute shaders.
  • Experience with low-level kernel optimization on mobile/resource-constrained devices.
  • Track record delivering measurable improvements in latency, throughput and memory.
  • Deep understanding of model-serving architectures and optimization techniques.
  • Experience writing GPU kernels for mobile devices.
  • End-to-end inference pipelines from model optimization to production on constrained hardware.
  • Ability to apply empirical research and systematic experimentation to solve bottlenecks.
  • Experience benchmarking inference systems across scales and environments.
  • Knowledge of distributed inference: tensor/pipeline/expert parallelism for large GPUs.
  • Understanding diffusion models and Vision Transformers architecture.
  • Familiarity with pruning, quantization, Flash Attention, KV Cache, speculative decoding.

Aufgaben

  • Design and deploy advanced model-serving architectures for high throughput and low latency.
  • Develop inference pipelines for mobile and edge platforms with constrained resources.
  • Set performance targets for latency, throughput, memory and reliability.
  • Build controlled benchmarks and collect latency, throughput and memory metrics.
  • Create datasets and simulations to evaluate model performance in real-world conditions.
  • Identify bottlenecks and implement solutions in batching, networking and memory management.
  • Develop custom GPU kernels and compute shaders for mobile hardware (MSL).
  • Apply inference optimization techniques including pruning, quantization and Flash Attention.
  • Design distributed inference with tensor/pipeline/expert parallelism for large GPU workloads.
  • Collaborate with cross-functional teams to integrate optimized inference into production and edge apps.
  • Document results and refine optimization strategies based on empirical evidence.
  • Monitor production performance and seek improvements in scalability and efficiency.

Kenntnisse

Latency optimization
Benchmarking
Empirical research
English communication
Tensor parallelism
Distributed inference

Ausbildung

PhD in NLP or ML
CS degree

Tools

Metal Shading Language (MSL)

Jobbeschreibung

Lever, Inc. is seeking an AI Research Engineer (Kernel & Inference Optimization) based in Switzerland. You will push the boundaries of efficient AI systems across hardware environments, balancing latency, throughput, and memory in model-serving architectures.

You will research and implement advanced inference strategies, develop custom GPU kernels, and collaborate across teams to translate research into production-ready infrastructure for edge devices.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Research Engineer (Kernel & Inference Optimization)
AI Research Engineer (Kernel & Inference Optimization)

Lever, Inc. • Schweiz

Remote
EUR 137.000 - 222.000
Remote-first team
International team
Cutting-edge AI research
+3
Senior AI Inference & HPC Systems Engineer
Senior AI Inference & HPC Systems Engineer

NVIDIA • Schweiz

Vor Ort
CHF 120.000 - 180.000
Senior ML Engineer, Inference & Optimization
Senior ML Engineer, Inference & Optimization

Nebius Group • Zürich

Vor Ort
CHF 150.000 - 220.000
Competitive compensation
Career growth and learning机会
Flexibility and ownership
+3
Senior Embedded Compiler Engineer for Edge AI & DSP
Senior Embedded Compiler Engineer for Edge AI & DSP

Ignite Next GmbH • Zürich

Hybrid
CHF 120.000 - 160.000
High-Performance Software Engineer — Verified ML Inference Engine
High-Performance Software Engineer — Verified ML Inference Engine

École polytechnique fédérale de Lausanne, EPFL • Lausanne

Vor Ort
CHF 110.000 - 150.000
Frontier AI models access
Travel for collaboration
Professional development
+1
Senior AI Inference & HPC Systems Engineer
Senior AI Inference & HPC Systems Engineer

NVIDIA • Zürich

Vor Ort
CHF 190.000 - 260.000
Lead Embedded Vision Systems Engineer (Real-Time AI Edge)
Lead Embedded Vision Systems Engineer (Real-Time AI Edge)

Luxoft • Schweiz

Remote
CHF 120.000 - 180.000
Senior HPC AI Network Architect: Scalable Infra
Senior HPC AI Network Architect: Scalable Infra

NVIDIA Corporation • Zürich

Vor Ort
CHF 180.000 - 240.000
Staff ML Engineer: Real-Time Multimodal Inference (Remote)
Staff ML Engineer: Real-Time Multimodal Inference (Remote)

Inworld AI • Schweiz

Vor Ort
CHF 100.000 - 140.000
Senior HPC & AI Network Architect for Scalable AI Infra
Senior HPC & AI Network Architect for Scalable AI Infra

NVIDIA • Zürich

Vor Ort
CHF 180.000 - 240.000