GPU Engineer

Ensimag Alumni

Paris

Sur place

EUR 50 000 - 70 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

Ensimag Alumni is seeking a talented engineer in Paris to work on GPU optimization for cutting-edge LLM inference at Kog. You'll engage in writing and profiling GPU kernels, optimizing computational paths, and contributing to a monokernel pipeline.

The ideal candidate has experience with CUDA or similar environments and a strong background from a top engineering school or relevant PhD. Join a dynamic team and push the boundaries of GPU engineering.

Qualifications

  • Experience in writing GPU kernels with performance constraints.
  • Understanding of hardware and GPU architecture.
  • Familiarity with inline PTX or CDNA ISA.

Responsabilités

  • Understand GPU internals and optimize computational sections.
  • Contribute to a monokernel pipeline for LLM inference.
  • Build profiling infrastructure for GPU programs.

Connaissances

GPU optimization
CUDA programming
PyTorch custom ops
Kernel optimization

Formation

Top engineering school or PhD with GPU work

Description du poste

Location

Paris

Contract
  • CDI
Compensation
  • A négocier
Company

Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding). The hot path is a monokernel implemented with handwritten CUDA (with PTX inline assembly) on NVIDIA, and HIP (with CDNA ISA inline assembly) on AMD. We optimize at the low level with engine/kernel/model co‑design, using reverse engineering to understand and exploit the details of how the GPU hardware works at the micro level. We are a team of 11 people, including 10 engineers and 4 PhDs. Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

Job Description

You will perform experiments to understand GPU internals, find creative solutions to accelerate critical computational sections used in LLM inference, and write optimized GPU kernels accordingly. Then test, profile, and optimize again.

Contribute to our monokernel pipeline, the single persistent GPU program that covers the full decode pass from QKV projection to LM head sampling, across AMD and NVIDIA architectures.

Work on low-level GPU optimization, including impossibly-fast grid synchronizations and inter‑GPU collectives, and optimized GEMM and attention kernels for specific batch sizes and context lengths.

Build profiling infrastructure inside a monokernel, including custom instrumentation, device‑timestamp frameworks, and per‑stage analysis to translate machine behavior into concrete engineering decisions.

Scale the stack to third‑party MoE models such as DeepSeek v4 and Qwen 3 to push generation speed on the models that matter in production today.

Contribute to building AI agents that will perform GPU Engineering research and kernel optimization autonomously, calibrated to hardware target and workload, starting from the inference foundations we are building now.

Qualifications
  • You have written GPU kernels where performance was the central constraint. Showing the code is a requirement to move forward in the process.
  • PyTorch custom ops are an acceptable starting point if the kernels show a genuine understanding of the hardware below the framework level.
  • Stronger signals include inline PTX or CDNA ISA in public repositories, experience with latency‑sensitive execution paths, understanding of why MBU matters more than MFU at batch size 1, and a background in inference engine components.
  • A top engineering school or a PhD with concrete GPU work counts, even without industry experience.
How to Apply

https://jobs.ashbyhq.com/kog/e3950334-a2a6-43cc-a744-df6c38683166

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Research Engineer
Research Engineer

Kog • Paris

Hybride
EUR 60 000 - 90 000
Direct access to AMD and NVIDIA GPUs
Creative and decision-influencing environment
Remote-friendly working model
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Ensimag Alumni • Paris

Sur place
EUR 50 000 - 70 000
Research Engineer (LLM Architecture)
Research Engineer (LLM Architecture)

Ensimag Alumni • Paris

Sur place
EUR 50 000 - 80 000
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Sur place
EUR 120 000 - 180 000
Stock options
Healthcare coverage
Pension contributions
+2
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago Inc. • Paris

Sur place
EUR 110 000 - 170 000
Stock options
Health insurance
Pension contributions
+1
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA AI • Aillas

Sur place
EUR 120 000 - 180 000
ML Systems Engineer — Inference Acceleration
ML Systems Engineer — Inference Acceleration

Arago • Paris

Sur place
EUR 90 000 - 130 000
Stock options
Healthcare coverage
Pension contributions
+2
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA • France

Sur place
EUR 90 000 - 150 000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • France

Sur place
EUR 110 000 - 160 000
Kokkos supporting for complex data discretization and unstructured meshes
Kokkos supporting for complex data discretization and unstructured meshes

Inria • Palaiseau

Sur place
EUR 40 000 - 65 000
Partial reimbursement of public transport costs
7 weeks of annual leave
Possibility of teleworking after 6 months
+2