Senior ML Engineer, LLM/VLM Inference Optimizer

Jobgether SRL

France

Sur place

EUR 90 000 - 150 000

Plein temps

Il y a 6 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Obtenez une réponse de cet employeur — un CV et une lettre de motivation adaptés exactement à ce qu’on recherche pour ce poste.

Passez les filtres ATS

Avantages offerts par ce poste

Competitive compensation
Career growth
Ownership over technical work
Collaborative environment
International teams

Résumé du poste

Jobgether SRL, partnering with a France-based team, seeks a Senior Machine Learning Engineer focused on LLM and VLM inference optimization. You will drive latency, throughput, and memory improvements across model families, engines, and deployment backends, coordinating with kernel, platform, and product teams.

Role involves benchmarking, safe rollouts, and collaboration with international AI teams, delivering measurable performance gains in production systems.

Qualifications

  • Strong Python and PyTorch software engineering skills.
  • Hands-on experience deploying/optimizing high-throughput transformer inference systems.
  • Experience with at least one modern inference stack (vLLM, TensorRT-LLM, Triton, etc.).
  • Ability to reason quantitatively about latency, throughput, cost, and resource use.

Responsabilités

  • Own optimization initiatives for model families and endpoints.
  • Evaluate inference engines and serving configurations.
  • Diagnose performance regressions during production rollouts.
  • Deploy, benchmark, and extend inference engines (vLLM, Triton, etc.).
  • Develop reproducible benchmark harnesses (TTFT, tokens/sec per GPU).
  • Collaborate with kernel, platform, and product teams.
  • Document design, performance reports, and rollout plans.

Connaissances

Python
PyTorch
LLM inference
Performance optimization
System benchmarking
CUDA/Triton
Communication
Open source

Outils

vLLM
TensorRT-LLM
Triton Inference Server
NVIDIA Dynamo
Ray Serve
KServe
SGLang
TorchServe

Description du poste

Jobgether SRL, partnering with a France-based team, seeks a Senior Machine Learning Engineer focused on LLM and VLM inference optimization. You will drive latency, throughput, and memory improvements across model families, engines, and deployment backends, coordinating with kernel, platform, and product teams.

Role involves benchmarking, safe rollouts, and collaboration with international AI teams, delivering measurable performance gains in production systems.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior ML Engineer: LLM Inference Optimization
Senior ML Engineer: LLM Inference Optimization

Lever, Inc. • France

Sur place
EUR 90 000 - 130 000
Competitive compensation
Career growth opportunities
Flexible work environment
+3
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Lever, Inc. • France

Sur place
EUR 90 000 - 130 000
Competitive compensation
Career growth opportunities
Flexible work environment
+3
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Jobgether SRL • France

Sur place
EUR 90 000 - 150 000
Competitive compensation
Career growth
Ownership over technical work
+2
Senior AI Inference Systems Engineer – GPU-Optimized ML
Senior AI Inference Systems Engineer – GPU-Optimized ML

NVIDIA • France

Hybride
EUR 120 000 - 190 000
Senior AI Infrastructure Engineer - LLM Platform Scale
Senior AI Infrastructure Engineer - LLM Platform Scale

Jobgether • France

Sur place
EUR 80 000 - 110 000
Competitive compensation package
Career growth and continuous learning
Flexibility and significant ownership
+3
Research Engineer — Multimodal AI & Agentic Systems
Research Engineer — Multimodal AI & Agentic Systems

H Company • Paris

Hybride
EUR 70 000 - 100 000
Competitive salary
Opportunities for professional growth
Collaborative and dynamic work environment
ML Inference Systems Engineer — High-Performance Accelerator
ML Inference Systems Engineer — High-Performance Accelerator

Arago Inc. • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Health insurance
Pension contributions
+1
Lead AI Engineer - Production ML & LLM Innovation
Lead AI Engineer - Production ML & LLM Innovation

Jobtailor • Paris

Sur place
EUR 90 000 - 120 000
ML Infrastructure Engineer for Scalable RL & Inference
ML Infrastructure Engineer for Scalable RL & Inference

White Circle • Paris

Hybride
EUR 157 701 - 306 642
Relocation package
Hybrid Paris/London work
Medical insurance
+2
ML Inference Systems Engineer - Accelerator Performance
ML Inference Systems Engineer - Accelerator Performance

Arago • Paris

Sur place
EUR 120 000 - 180 000
Stock options
Healthcare coverage
Pension contributions
+2