Senior LLMOps, RAG & infrastructure

Adlin Science

Paris

Sur place

EUR 90 000 - 130 000

Plein temps

Il y a 2 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Une candidature sur mesure pour ce poste — un CV personnalisé et une lettre de motivation qui correspondent directement à l’offre.

Passez les filtres ATS

Avantages offerts par ce poste

Restaurant d'entreprise
Tickets restauration

Résumé du poste

Adlin Science crée une plateforme de gouvernance et de gestion de données biomédicales multimodales et recherche un responsable LLM pour piloter l’architecture et l’intégration du socle IA au sein de l’équipe Data. Vous contribuerez à définir l’architecture modulaire et à déployer des solutions sans dépendance à des API externes en production.

Le rôle requiert une forte expérience en ML engineering/MLOps, une maîtrise avancée de Python et PyTorch, et une sensibilité à la sécurité et à la

Qualifications

  • Maîtrise de déploiement et exploitation de LLM/NLP sous contraintes de performance.
  • Maîtrise de Python, PyTorch et de l’écosystème Transformers.
  • Maîtrise d’au moins un moteur d’inférence haute performance et expérience avec quantization/GPU sizing.
  • Expérience pratique avec les architectures RAG, embeddings, recherche hybride et évaluation des systèmes génératifs.
  • Solides compétences Docker, Linux, CI/CD, observabilité et gestion des artefacts en environnements contrôlés.
  • Culture sécurité: gestion des secrets, contrôle d’accès, encryption et chaîne d’approvisionnement logicielle.
  • Capacité à documenter, expliquer les compromis et travailler avec plusieurs équipes sans dette technique isolée.
  • Connaissances Kubernetes, Terraform, MLflow, Prometheus/Grafana, OpenSearch et bases de données vectorielles locales comme Qdrant, Milvus, pgvector.
  • PEFT/LoRA/QLoRA, distillation, entraînement distribué et préparation de jeux de données supervisés.
  • Connaissances des contraintes du secteur de la santé, GDPR et environnements sécurisés.
  • Expérience en environnements isolés (air-gapped) et sur site, procédures d’importation contrôlées.
  • Contribution open-source ou évaluation critique de papiers scientifiques.
  • Autonomie décisionnelle et esprit critique pour avancer même en l’absence de définition complète.
  • Esprit proactif et communicatif, capable d’échanger avec des publics non techniques.

Responsabilités

  • Sélectionner, tester et qualifier des modèles ouverts compatibles déploiement local et contraintes licensing/confidentialité.
  • Définir une architecture modulaire permettant de changer le modèle ou la stratégie RAG sans couplage fort.
  • Mettre en œuvre le RAG local: embeddings, recherche vectorielle, hybrid search, reranking, filtrage et gestion des citations.
  • Déployer et faire opérer des moteurs d’inférence locaux tels que vLLM, llama.cpp, TGI, TensorRT-LLM, Triton ou équivalents.
  • Optimiser latence, débit, mémoire et stabilité via quantization, batching, KV-cache, parallélisme et choix de formats.
  • Mettre en place tests de charge, résilience et non-régression avant chaque release production.
  • Concevoir des pipelines packaging/versioning/validation/promotion/rollback des modèles, prompts et indices.
  • Définir un processus contrôlé d’importation de modèles et dépendances dans l’environnement sécurisé.
  • Containeriser les composants et les intégrer au CI/CD et outils d’orchestration.
  • Produire runbooks, procédures d’incident et documents d’architecture pour les revues sécurité et qualité.

Connaissances

LLM deployment
Python
PyTorch & Transformers
Inference engine
RAG architectures
Docker & CI/CD
Security culture
Documentation & communication
Kubernetes & Terraform
MLflow / Prometheus/Grafana
OpenSearch / vector DBs
PEFT/LoRA/QLoRA
GDPR & on-prem
Autonomie & proactivité
Open-source evaluation

Formation

Master’s degree or PhD in CS/AI

Outils

Qdrant
Milvus
pgvector
OpenSearch
Docker
Kubernetes
Terraform
MLflow
Prometheus/Grafana

Description du poste

Location: Paris

Type de contrat : Full-time permanent position (CDI)

Start date: As soon as possible

Compensation: Based on experience and profile

Adlin Science develops a platform for the governance, quality management and exploitation of multimodal biomedical data. We are strengthening the Data team with a senior profile capable of building the platform’s LLM/IA foundation.

The role is not limited to model serving. It covers usage architecture, selection and qualification of open models, local RAG, evaluation, inference optimization, observability and real-world operation on sensitive data

Role positioning
  • You are the LLM technical referent within the Data team and work closely with the Backend, DevOps, Security, Product and domain expert teams.
  • You design the components specific to the LLM chain and integrate them into the existing Adlin architecture.
  • You do not redefine infrastructure, backend or security standards on your own: you rely on the choices, tools and constraints defined with the responsible teams.
  • You prioritize installable, auditable and maintainable solutions, with no dependency on an external API in production.
Main Responsibilities
  • Select, test and qualify open models compatible with local deployment and licensing, confidentiality and redistribution constraints.
  • Define a modular architecture allowing the model, inference engine or RAG strategy to be changed without excessive coupling.
  • Implement local RAG: embeddings, lexical and vector search, hybrid search, reranking, metadata filters and citation management.
  • Deploy and operate local inference engines such as vLLM, llama.cpp, TGI, TensorRT-LLM, Triton or equivalent solutions depending on the need
  • Optimize latency, throughput, memory consumption and stability through quantization, continuous batching, KV-cache management, parallelism and appropriate format choices.
  • Set up load, resilience and non-regression tests before each production release.
  • Build pipelines for packaging, versioning, validation, promotion and rollback of models, prompts, embeddings, indexes and configurations
  • Define a controlled process for importing models and dependencies into the secure environment: signed artifacts, integrity checks, inventory, vulnerability scans and approval procedure.
  • Containerize components and integrate them with the CI/CD and orchestration tools selected by the Backend / DevOps teams
  • Produce runbooks, incident procedures, architecture files, test evidence and the elements required for security and quality reviews.
Required Skills
  • Hands-on experience deploying and operating LLMs or NLP/ML systems under strong performance constraints.
  • Excellent command of Python, PyTorch, the Transformers ecosystem and the architectural principles of language models.
  • Mastery of at least one high-performance inference engine and experience with quantization and GPU sizing.
  • Practical experience with RAG architectures, embeddings, hybrid search, reranking and evaluation of generative systems.
  • Strong Docker, Linux, CI/CD, observability and artifact management skills in controlled environments.
  • Security culture: secrets, access control, encryption, isolation, software supply chain and sensitive data processing.
  • Ability to document, explain trade-offs and work with multiple teams without creating isolated technical debt.
  • Nice to have Kubernetes, Terraform, MLflow, Prometheus/Grafana, OpenSearch and locally deployable vector databases such as Qdrant, Milvus or pgvector.
  • PEFT/LoRA/QLoRA fine-tuning, distillation, distributed training and preparation of supervised datasets.
  • Knowledge of healthcare sector constraints, GDPR, sensitive data and regulated environments.
  • Experience with air-gapped, on-premise, appliance or edge environments, including controlled import procedures.
  • Open-source contribution or experience reading and critically evaluating scientific papers.
  • FunctionalAutonomy and decision-making: you know how to scope your work and move forward even when things are not fully defined.
  • Proactive mindset and critical thinking.
  • Comfortable in dynamic, evolving environments: you thrive when processes are still being shaped.
  • Open and constructive communication: you can break down technical concepts for non-technical audiences, share knowledge freely and foster collaborative dialogue.
Profile
  • Master’s degree or PhD in computer science, AI, machine learning or a related field, or equivalent experience.
  • Significant experience in ML Engineering, MLOps or AI infrastructure, with at least one LLM deployment actually operated in production.
  • Hands-on profile, able to prototype, code, benchmark, diagnose and industrialize.
  • Rigor, autonomy, team spirit and ability to work in a context where traceability and quality come first.
  • Fluent technical English.
Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Design, develop & deploy ML/LLM models, data pipelines & agentic AI solutions for int/ext projects
Design, develop & deploy ML/LLM models, data pipelines & agentic AI solutions for int/ext projects

OUTSCALE • Saint-Cloud

Sur place
EUR 70 000 - 110 000
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Lever, Inc. • France

Sur place
EUR 90 000 - 130 000
Competitive compensation
Career growth opportunities
Flexible work environment
+3
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Jobgether SRL • France

Sur place
EUR 90 000 - 150 000
Competitive compensation
Career growth
Ownership over technical work
+2
Lead LLM Engineer
Lead LLM Engineer

Leonar • Paris

Sur place
EUR 90 000 - 150 000
Senior LLMOps, RAG & Infra Architect
Senior LLMOps, RAG & Infra Architect

Adlin Science • Paris

Sur place
EUR 90 000 - 130 000
Restaurant d'entreprise
Tickets restauration
Research Engineer – AI Agents & LLM Systems F/M
Research Engineer – AI Agents & LLM Systems F/M

Adoc Talent Management • Paris

Sur place
EUR 85 000 - 120 000
LLM Engineer
LLM Engineer

KDCI • Job

Hybride
USD 120 000 - 190 000
Research Engineer – AI Agents & LLM Systems F/M
Research Engineer – AI Agents & LLM Systems F/M

Adoc Tm • Paris

Sur place
EUR 90 000 - 120 000
data scientist / llm engineer senior (generative ai) - (h/f).
data scientist / llm engineer senior (generative ai) - (h/f).

MissionHandicap • Toulouse

Sur place
EUR 60 000 - 100 000
ML Infrastructure Engineer
ML Infrastructure Engineer

White Circle • Paris

Sur place
EUR 157 701 - 306 642
Relocation package
Hybrid Paris/London work
Medical insurance
+2