Research Engineer

Kog

Paris

Sur place

EUR 60 000 - 90 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Direct access to AMD and NVIDIA GPUs
Creative and decision-influencing environment
Remote-friendly working model

Résumé du poste

Kog is seeking an innovative architect to design and optimize model architectures focused on inference behavior. You'll experiment with cutting-edge AI technologies in a dynamic team setting.

In this role, you'll influence our model development processes while conducting research that enhances execution speed. A collaborative environment allows your technical judgment to shape key decisions with direct impact on system evolution.

Qualifications

  • Experience adapting or modifying existing model architectures.
  • Understanding of communication structure and layer dependencies affecting inference behavior.
  • Fluency in Transformers and MoE with depth to reason across trade-offs.

Responsabilités

  • Design new model architecture variants with execution constraints as input.
  • Explore inference‑aware architectural variants and discover scalable compounds.
  • Own the post-training pipeline for open-weight models optimized for inference speed.
  • Scale large MoE models and analyze communication patterns at inference.
  • Write and present findings in research papers at top venues.

Connaissances

Model design
Architectural optimization
Post-training methods
AI problem-solving

Formation

PhD or equivalent in a relevant field

Description du poste

About Kog

Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding).

We co-design the model architecture and the execution engine together. Our Laneformer model uses Delayed Tensor Parallelism (DTP), a novel architecture that restructures the Transformer dependency graph so inter-GPU communication overlaps with computation rather than blocking it.

We pretrained a 2B-parameter DTP model on 6T tokens on 256 H100 GPUs.

We are a team of 11 people, including 10 engineers and 4 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

What you will work on

You will imagine, design and run experiments to understand how architectural decisions propagate through inference behavior, morph existing open-weight models into architecture variants optimized for speed, and turn findings into measurable gains in generation speed and model quality.

  • Design new model architecture variants, including routing strategies, attention mechanisms, and MoE structure, with execution constraints as a first-order design input.
  • Extend the Laneformer thesis by exploring inference‑aware architectural variants such as DTP, Ladder Residual, and PT‑Transformer, and finding what compounds at scale.
  • Own the post‑training pipeline across fine‑tuning, evaluation methodology, and adaptation of existing open‑weight models toward architecture variants optimized for inference speed.
  • Scale the stack to large MoE models such as DeepSeek v4 and Qwen 3, working through routing, expert parallelism, and communication patterns at inference time.
  • Write up findings as research papers, submit them to top venues, and present them at conferences.
What we look for

You are rigorous, curious, and comfortable working at the intersection of model design and hardware constraints.

You have worked on complex AI problems and have something concrete to show for it. A paper, a repository, a thesis, or a side project with evidence of serious technical thinking is what we want to see.

Strong signals include experience adapting or modifying existing model architectures, understanding of how communication structure and layer dependencies affect inference behavior, and fluency in Transformers and MoE with enough depth to reason across trade‑offs.

Experience in post‑training methods such as fine‑tuning, preference optimization, or quantization is a plus, even without production‑scale exposure.

What we offer
  • Direct access to AMD and NVIDIA datacenter GPUs from day one
  • A team where creativity and technical judgment carry weight and where the people closest to the problem shape the key decisions
  • Problems that sit on the critical path of model execution speed and that directly influence what the system can become
  • A remote‑friendly working model, though you'll spend at least 50% of your time in our Paris office
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

GPU Engineer
GPU Engineer

Ensimag Alumni • Paris

Sur place
EUR 50 000 - 70 000
Research Engineer (LLM Architecture)
Research Engineer (LLM Architecture)

Ensimag Alumni • Paris

Sur place
EUR 50 000 - 80 000
Research Engineer: Inference Architect for Fast LLMs Remote
Research Engineer: Inference Architect for Fast LLMs Remote

Kog • Paris

Hybride
EUR 60 000 - 90 000
Direct access to AMD and NVIDIA GPUs
Creative and decision-influencing environment
Remote-friendly working model
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Sur place
EUR 120 000 - 180 000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago Inc. • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Health insurance
Pension contributions
+1
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Sur place
EUR 90 000 - 130 000
Stock options
Healthcare coverage
Pension contributions
+2
Research Scientist / Research Engineer
Research Scientist / Research Engineer

adaption • Paris

Hybride
EUR 90 000 - 130 000
Flexible work
Travel stipend
Lunch stipend
+1
Research Engineer, Model Inference & Serving - Paris
Research Engineer, Model Inference & Serving - Paris

H Company • Paris

Sur place
EUR 60 000 - 90 000
Competitive salary
Opportunities for professional growth
Collaborative multicultural team
MLOps Engineer
MLOps Engineer

White Circle • Paris

Sur place
EUR 65 000 - 85 000
Paid time off
Comprehensive medical insurance
Team off-sites twice a year
ML Runtime Engineer
ML Runtime Engineer

ATX Venture Partners • Occitanie

Sur place
EUR 85 000 - 110 000
Equity
Paid time off
Health insurance
+4