Location: Paris
Type de contrat : Full-time permanent position (CDI)
Start date: As soon as possible
Compensation: Based on experience and profile
Adlin Science develops a platform for the governance, quality management and exploitation of multimodal biomedical data. We are strengthening the Data team with a senior profile capable of building the platform’s LLM/IA foundation.
The role is not limited to model serving. It covers usage architecture, selection and qualification of open models, local RAG, evaluation, inference optimization, observability and real-world operation on sensitive data
Role positioning
- You are the LLM technical referent within the Data team and work closely with the Backend, DevOps, Security, Product and domain expert teams.
- You design the components specific to the LLM chain and integrate them into the existing Adlin architecture.
- You do not redefine infrastructure, backend or security standards on your own: you rely on the choices, tools and constraints defined with the responsible teams.
- You prioritize installable, auditable and maintainable solutions, with no dependency on an external API in production.
Main Responsibilities
- Select, test and qualify open models compatible with local deployment and licensing, confidentiality and redistribution constraints.
- Define a modular architecture allowing the model, inference engine or RAG strategy to be changed without excessive coupling.
- Implement local RAG: embeddings, lexical and vector search, hybrid search, reranking, metadata filters and citation management.
- Deploy and operate local inference engines such as vLLM, llama.cpp, TGI, TensorRT-LLM, Triton or equivalent solutions depending on the need
- Optimize latency, throughput, memory consumption and stability through quantization, continuous batching, KV-cache management, parallelism and appropriate format choices.
- Set up load, resilience and non-regression tests before each production release.
- Build pipelines for packaging, versioning, validation, promotion and rollback of models, prompts, embeddings, indexes and configurations
- Define a controlled process for importing models and dependencies into the secure environment: signed artifacts, integrity checks, inventory, vulnerability scans and approval procedure.
- Containerize components and integrate them with the CI/CD and orchestration tools selected by the Backend / DevOps teams
- Produce runbooks, incident procedures, architecture files, test evidence and the elements required for security and quality reviews.
Required Skills
- Hands-on experience deploying and operating LLMs or NLP/ML systems under strong performance constraints.
- Excellent command of Python, PyTorch, the Transformers ecosystem and the architectural principles of language models.
- Mastery of at least one high-performance inference engine and experience with quantization and GPU sizing.
- Practical experience with RAG architectures, embeddings, hybrid search, reranking and evaluation of generative systems.
- Strong Docker, Linux, CI/CD, observability and artifact management skills in controlled environments.
- Security culture: secrets, access control, encryption, isolation, software supply chain and sensitive data processing.
- Ability to document, explain trade-offs and work with multiple teams without creating isolated technical debt.
- Nice to have Kubernetes, Terraform, MLflow, Prometheus/Grafana, OpenSearch and locally deployable vector databases such as Qdrant, Milvus or pgvector.
- PEFT/LoRA/QLoRA fine-tuning, distillation, distributed training and preparation of supervised datasets.
- Knowledge of healthcare sector constraints, GDPR, sensitive data and regulated environments.
- Experience with air-gapped, on-premise, appliance or edge environments, including controlled import procedures.
- Open-source contribution or experience reading and critically evaluating scientific papers.
- FunctionalAutonomy and decision-making: you know how to scope your work and move forward even when things are not fully defined.
- Proactive mindset and critical thinking.
- Comfortable in dynamic, evolving environments: you thrive when processes are still being shaped.
- Open and constructive communication: you can break down technical concepts for non-technical audiences, share knowledge freely and foster collaborative dialogue.
Profile
- Master’s degree or PhD in computer science, AI, machine learning or a related field, or equivalent experience.
- Significant experience in ML Engineering, MLOps or AI infrastructure, with at least one LLM deployment actually operated in production.
- Hands-on profile, able to prototype, code, benchmark, diagnose and industrialize.
- Rigor, autonomy, team spirit and ability to work in a context where traceability and quality come first.
- Fluent technical English.