Senior AI Inference Architect for Large-Scale Deployments

Talanto

France

Hybride

EUR 68 000 - 150 000

Plein temps

Il y a 44 heures
Soyez parmi les premiers à postuler
Générateur de candidature

Une candidature complète en une minute — un CV et une lettre de motivation personnalisés, prêts à être envoyés.

Passez les filtres ATS

Résumé du poste

Talanto is seeking a Senior Solutions Architect to lead AI inference strategy across EMEA. You will guide engagements from proof of concept to production-scale deployments, architect high-performance inference pipelines, and optimize GPU utilization for large-scale AI workloads.

The role requires deep expertise in LLM/VLM inference, transformer acceleration, and Kubernetes-based orchestration. Join a frontier AI environment with strong collaboration across research and enterprise teams.

Qualifications

  • MS or PhD in CS, Engineering, or equivalent.
  • 8+ years in AI/ML infrastructure and production deployment at scale.
  • Deep expertise in transformer inference acceleration: quantization, speculative decoding, and KV cache optimization.
  • Understanding GPU memory hierarchies and low-latency networking for inference.
  • Proven track record leading technical initiatives.
  • Excellent communication with researchers, engineers, and executives.

Responsabilités

  • Define the technical direction for AI inference across EMEA.
  • Lead inference strategy from PoC to production deployments.
  • Architect and optimize high-performance inference pipelines using listed backends.
  • Translate customer insights into actionable feedback shaping stack roadmaps.

Connaissances

AI/ML infrastructure
Leadership
Technical collaboration
Transformer inference

Formation

MS/PhD in CS or Engineering

Outils

Dynamo
TensorRT-LLM
vLLM
SGLang
Kubernetes
NIM

Description du poste

Talanto is seeking a Senior Solutions Architect to lead AI inference strategy across EMEA. You will guide engagements from proof of concept to production-scale deployments, architect high-performance inference pipelines, and optimize GPU utilization for large-scale AI workloads.

The role requires deep expertise in LLM/VLM inference, transformer acceleration, and Kubernetes-based orchestration. Join a frontier AI environment with strong collaboration across research and enterprise teams.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior AI Inference Architect - Large-Scale GPU Pipelines
Senior AI Inference Architect - Large-Scale GPU Pipelines

NVIDIA • France

Sur place
EUR 150 000 - 210 000
Senior AI Inference Architect for Multi-Node GPU Scale
Senior AI Inference Architect for Multi-Node GPU Scale

NVIDIA Corporation • France

Hybride
EUR 120 000 - 180 000
Senior AI Solutions Architect - Vision & Diffusion
Senior AI Solutions Architect - Vision & Diffusion

NVIDIA France • France

Sur place
EUR 120 000 - 180 000
Senior Solutions Architect – Large Scale AI Inference
Senior Solutions Architect – Large Scale AI Inference

NVIDIA Corporation • France

Hybride
EUR 120 000 - 180 000
Enterprise AI & LLM Solutions Architect
Enterprise AI & LLM Solutions Architect

Tencent • France

Sur place
EUR 90 000 - 130 000
Senior Solutions Architect – Large Scale AI Inference
Senior Solutions Architect – Large Scale AI Inference

NVIDIA • France

Sur place
EUR 150 000 - 210 000
Senior Data Infra Engineer for Scalable AI Compute
Senior Data Infra Engineer for Scalable AI Compute

Mistral.ai • Paris

Hybride
EUR 90 000 - 140 000
Senior AI Deployment Strategist - Enterprise AI Leader
Senior AI Deployment Strategist - Enterprise AI Leader

Mistral • Paris

Sur place
EUR 140 000 - 200 000
Competitive cash salary
Equity
Daily lunch vouchers
Senior Solutions Architect – Large Scale Neural Networks Inference
Senior Solutions Architect – Large Scale Neural Networks Inference

Talanto • France

Hybride
EUR 68 000 - 150 000
Enterprise AI & LLM Solutions Architect
Enterprise AI & LLM Solutions Architect

Tencent • Paris

Sur place
EUR 110 000 - 140 000