Senior ML Engineer: LLM Inference Optimization

Lever, Inc.

France

Sur place

EUR 90 000 - 130 000

Plein temps

Il y a 7 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Une candidature complète en une minute — CV personnalisé et lettre de motivation, prêts à envoyer.

Passez les filtres ATS

Avantages offerts par ce poste

Competitive compensation
Career growth opportunities
Flexible work environment
Ownership over technical work
International team
Impactful AI infrastructure projects

Résumé du poste

Lever, Inc. seeks a Senior Machine Learning Engineer to optimize LLM and VLM inference in a France-based role. You will drive optimizations from model artifacts through production, collaborating across kernel, platform, and research teams to reduce latency, increase throughput, and lower token cost.

The role demands hands-on experience with modern inference stacks and quantization-aware training, with a focus on measurable production improvements and robust benchmarking.

Qualifications

  • Strong software engineering in Python and PyTorch.
  • Hands-on experience deploying, operating, or optimizing LLM/VLM inference systems.
  • Experience with modern inference stacks such as vLLM, SGLang, TensorRT-LLM, Triton, NVIDIA Dynamo, Ray Serve, KServe.
  • Understanding of transformer inference bottlenecks: KV cache, attention, memory bandwidth, batching, parallelism, long-context serving.
  • Ability to reason about latency, throughput, quality, resource use, and cost trade-offs.
  • Experience diagnosing performance problems and translating findings into production improvements.
  • Strong communication and collaboration with research, kernel, infrastructure, product, and customer teams.
  • Experience with quantization-aware training, post-training quantization, FP8/INT8/INT4, and related optimizations is a plus.
  • Familiarity with distillation, speculative decoding, multi-token prediction, or acceleration approaches is advantageous.

Responsabilités

  • Own optimization initiatives for specific model families, customer endpoints, and inference serving backends.
  • Evaluate inference engines and recommend practical serving configurations based on workload requirements.
  • Diagnose and resolve model quality, performance, and reliability regressions during production rollouts.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token.
  • Deploy, configure, benchmark, and extend modern inference engines like vLLM, SGLang, TensorRT-LLM, Triton, NVIDIA Dynamo, or equivalents.
  • Build and productionize model-compression workflows including quantization and distillation.
  • Implement advanced inference techniques like speculative decoding, KV-cache optimization, and disaggregated prefill/decode serving.
  • Develop reproducible benchmark harnesses covering TTFT, tokens/sec per GPU, latency p95/p99, memory usage, and cost per token.
  • Partner with GPU kernel and platform engineers to identify bottlenecks across model code and infrastructure.
  • Investigate performance trade-offs and use benchmarks to guide optimizations.
  • Produce design docs, performance reports, and rollout plans for internal and customer stakeholders.
  • Contribute to safe, measurable production rollouts of inference improvements.

Connaissances

Python
PyTorch
Performance optimization
Quantization
Quantization-aware training
Inference systems
Benchmarking
Communication skills

Outils

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
NVIDIA Dynamo
Ray Serve
KServe
CUDA

Description du poste

Lever, Inc. seeks a Senior Machine Learning Engineer to optimize LLM and VLM inference in a France-based role. You will drive optimizations from model artifacts through production, collaborating across kernel, platform, and research teams to reduce latency, increase throughput, and lower token cost.

The role demands hands-on experience with modern inference stacks and quantization-aware training, with a focus on measurable production improvements and robust benchmarking.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior ML Engineer, LLM/VLM Inference Optimizer
Senior ML Engineer, LLM/VLM Inference Optimizer

Jobgether SRL • France

Sur place
EUR 90 000 - 150 000
Competitive compensation
Career growth
Ownership over technical work
+2
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Jobgether SRL • France

Sur place
EUR 90 000 - 150 000
Competitive compensation
Career growth
Ownership over technical work
+2
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Lever, Inc. • France

Sur place
EUR 90 000 - 130 000
Competitive compensation
Career growth opportunities
Flexible work environment
+3
Lead AI Engineer - Production ML & LLM Innovation
Lead AI Engineer - Production ML & LLM Innovation

Jobtailor • Paris

Sur place
EUR 90 000 - 120 000
ML Inference Systems Engineer — High-Performance Accelerator
ML Inference Systems Engineer — High-Performance Accelerator

Arago Inc. • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Health insurance
Pension contributions
+1
Senior AI Inference Systems Engineer – GPU-Optimized ML
Senior AI Inference Systems Engineer – GPU-Optimized ML

NVIDIA • France

Hybride
EUR 120 000 - 190 000
ML Infrastructure Engineer for Scalable RL & Inference
ML Infrastructure Engineer for Scalable RL & Inference

White Circle • Paris

Hybride
EUR 157 701 - 306 642
Relocation package
Hybrid Paris/London work
Medical insurance
+2
ML Engineer, NLP & LLMs | Remote‑Friendly + Stock Options
ML Engineer, NLP & LLMs | Remote‑Friendly + Stock Options

Dormont Manufacturing Co • France

Hybride
EUR 55 000 - 85 000
Competitive package
Stock options
Free health insurance
+5
ML Research Engineer — LLM Training & Safety
ML Research Engineer — LLM Training & Safety

White Circle • Paris

Hybride
EUR 105 097 - 218 952
Paid time off
Comprehensive medical insurance
Team off-sites twice a year
Research Engineer — Multimodal AI & Agentic Systems
Research Engineer — Multimodal AI & Agentic Systems

H Company • Paris

Hybride
EUR 70 000 - 100 000
Competitive salary
Opportunities for professional growth
Collaborative and dynamic work environment