Research Engineer (LLM Architecture)

Ensimag Alumni

Paris

Sur place

EUR 50 000 - 80 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

Kog recherche un ingénieur en architecture IA pour concevoir et optimiser des modèles d'inférence. Votre rôle clé sera d'imaginer et réaliser des expériences qui amélioreront la vitesse et la qualité des modèles existants.

Vous travaillerez sur des architectures innovantes, en tenant compte des contraintes matérielles, et contribuerez à la recherche en soumettant vos découvertes à des conférences. Des compétences dans l'adaptation d'architectures et une expérience en fine-tuning sont souhaitées.

Qualifications

  • Vous avez travaillé sur des problèmes complexes d'IA avec des résultats concrets.
  • Expérience dans l'adaptation ou la modification d'architectures de modèles existantes.
  • Compréhension approfondie des architectures Transformers et MoE.

Responsabilités

  • Imaginer, concevoir et réaliser des expériences.
  • Concevoir de nouvelles variantes d'architecture de modèles.
  • Étendre la thèse Laneformer par des variantes architecturales.

Connaissances

Adaptation des architectures de modèles
Compréhension des structures de communication
Expertise en Transformers et MoE

Description du poste

Postée le 19 juin

Lieu : Paris

  • Contrat : CDI
  • Rémunération : A négocier

Société : Kog

Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding).

We co‑design the model architecture and the execution engine together. Our Laneformer model uses Delayed Tensor Parallelism (DTP), a novel architecture that restructures the Transformer dependency graph so inter‑GPU communication overlaps with computation rather than blocking it.

We pretrained a 2B‑parameter DTP model on 6T tokens on 256 H100 GPUs.

We are a team of 11 people, including 10 engineers and 4 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

Description du poste

You will imagine, design and run experiments to understand how architectural decisions propagate through inference behavior, morph existing open‑weight models into architecture variants optimized for speed, and turn findings into measurable gains in generation speed and model quality.

Design new model architecture variants, including routing strategies, attention mechanisms, and MoE structure, with execution constraints as a first‑order design input.

Extend the Laneformer thesis by exploring inference‑aware architectural variants such as DTP, Ladder Residual, and PT‑Transformer, and finding what compounds at scale.

Own the post‑training pipeline across fine‑tuning, evaluation methodology, and adaptation of existing open‑weight models toward architecture variants optimized for inference speed.

Scale the stack to large MoE models such as DeepSeek v4 and Qwen 3, working through routing, expert parallelism, and communication patterns at inference time.

Write up findings as research papers, submit them to top venues, and present them at conferences.

Contribute to building AI agents that will perform architecture research and training experiments autonomously, starting from the research foundations we are building now.

Profil recherché

You are rigorous, curious, and comfortable working at the intersection of model design and hardware constraints.

You have worked on complex AI problems and have something concrete to show for it. A paper, a repository, a thesis, or a side project with evidence of serious technical thinking is what we want to see.

Strong signals include experience adapting or modifying existing model architectures, understanding of how communication structure and layer dependencies affect inference behavior, and fluency in Transformers and MoE with enough depth to reason across trade‑offs.

Experience in post‑training methods such as fine‑tuning, preference optimization, or quantization is a plus, even without production‑scale exposure.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Research Engineer
Research Engineer

Kog • Paris

Hybride
EUR 60 000 - 90 000
Direct access to AMD and NVIDIA GPUs
Creative and decision-influencing environment
Remote-friendly working model
GPU Engineer
GPU Engineer

Ensimag Alumni • Paris

Sur place
EUR 50 000 - 70 000
LLM Architecture Research Engineer
LLM Architecture Research Engineer

Ensimag Alumni • Paris

Sur place
EUR 50 000 - 80 000
Data Scientist (H/F)
Data Scientist (H/F)

France Travail • Paris

Sur place
Complémentaire santé
Titres restaurant
Indemnité transports
+1
LLM Engineer (F/H) – CDI – Annecy
LLM Engineer (F/H) – CDI – Annecy

Datsup • Annecy

Sur place
EUR 60 000 - 90 000
LLM Engineer
LLM Engineer

CLEEVEN • Auvergne-Rhône-Alpes

Sur place
EUR 85 000 - 110 000
Senior ML Engineer - NLP/LLM Specialist - Bordeaux (H/F)
Senior ML Engineer - NLP/LLM Specialist - Bordeaux (H/F)

Kicklox • Bordeaux

Hybride
Collaboration avec des leaders d'industrie
Possibilité de CDI avec télétravail
Projets novateurs en IA
Développeur Fullstack LLM H/F - Paris
Développeur Fullstack LLM H/F - Paris

Webnet • Paris

Sur place
EUR <100 000
ML platform senior DevOps engineer Frelance
ML platform senior DevOps engineer Frelance

Cherry Pick • Paris

Sur place
EUR 75 000 - 100 000
Senior Applied ML Engineer (3D)
Senior Applied ML Engineer (3D)

Chat3D • Lyon

Sur place
EUR 90 000 - 110 000