Research Engineer, Model Inference & Serving - Paris

hcompany

Paris

Hybride

EUR 90 000 - 130 000

Plein temps

14 jours+
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV personnalisé et une lettre de motivation en environ une minute.

Passez les filtres ATS

Avantages offerts par ce poste

Hybrid work (3 days in office)
Professional growth
Multicultural team

Résumé du poste

H is hiring a Research Engineer to build and operate the inference stack for multimodal agentic models. You will optimize latency, throughput, and cost across the model serving pipeline, and co-design with the Models team on training-time decisions affecting inference.

Based in Paris or London, this hybrid role expects in-office presence three days a week with periodic travel between offices. You will collaborate with cross-functional teams to deploy production ML infrastructure using Kubernetes

Qualifications

  • Advanced degree with research output in AI, ML, or systems.
  • Strong software engineering background with production-grade experience.
  • Proficiency in Python and at least one systems language (Rust, C++, or Go).
  • Hands-on experience with deep learning frameworks (PyTorch, JAX).
  • Knowledge of distributed systems and modern ML infrastructure (Kubernetes).
  • Experience with multimodal architectures and transformer models.

Responsabilités

  • Build and operate the inference stack that serves multimodal agentic models.
  • Improve latency, throughput, and cost of model serving across the stack.
  • Research and implement inference techniques for agent workloads.
  • Co-design with Models team on training-time decisions affecting inference.
  • Collaborate with cross-functional teams to deploy inference in AI products.
  • Evaluate inference, serving, and hardware platforms; communicate findings.

Connaissances

Software engineering
Python proficiency
Systems language (Rust/C++/Go)
PyTorch
Distributed systems
Kubernetes
ML basics (transformers, multimodal)
Communication skills

Formation

Advanced degree with research output

Outils

Kubernetes
vLLM
SGLang
TensorRT-LLM
CUDA

Description du poste

Research Engineer, Model Inference & Serving

About H: H exists to push the boundaries of superintelligence with agentic AI. By automating complex, multi-step tasks typically performed by humans, AI agents will help unlock full human potential. H is hiring the world's best AI talent, seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute.

About the Team: The Inference team builds and operates the systems that serve H's foundational models in production. We focus on multimodal inference and serving for Computer Use Agents, optimizing across both the inference engine layer (e.g., vLLM, SGLang) and the model serving layer (e.g., disaggregated inference, intelligent routing). Agentic inference brings constraints around context length, multimodality, and tool calls, which we address by co-designing with the Models team on training-time choices and with the agent teams on how models are deployed. We operate at the intersection of research and production, translating cutting-edge inference techniques into the systems that power H's next generation of agents. We are looking for strong engineers excited about inference to join the team and help shape the systems behind superintelligent AI.

Key Responsibilities:
  • Build and operate the inference stack that serves H's multimodal agentic models

  • Improve latency, throughput, and cost of model serving across the stack

  • Research and implement inference techniques tailored to agent workloads

  • Co-design with the Models team on training-time decisions that affect inference

  • Collaborate with cross-functional teams to integrate inference into agentic AI products

  • Evaluate inference, serving, and hardware platforms, and communicate findings to stakeholders

  • Stay current with advancements in inference, model serving, and accelerator technology

Requirements:
  • Technical skills:

    • Strong software engineering track record

    • Proficient in Python and at least one systems language (Rust, C++, or Go)

    • Hands-on experience with deep learning frameworks (PyTorch, JAX), preferably in an industry setting

    • Solid distributed systems fundamentals

    • Experience working in a modern cloud environment and with production ML infrastructure (Kubernetes, etc.)

    • Working knowledge of modern ML, including transformers and multimodal architectures

  • Research skills:

    • Research engagement: an advanced degree with research output, or publications at top-tier AI or systems venues (e.g., NeurIPS, ICML, MLSys, OSDI), research internships, or substantive open-source contributions

  • Soft skills:

    • Excellent communication and presentation skills

    • Strong collaboration and teamwork skills

    • Passion for inference and AI

  • Preferred qualifications:

    • Startup experience

    • Hands-on experience with inference frameworks (vLLM, SGLang, TensorRT-LLM)

    • Writing or modifying GPU kernels (CUDA, Triton, etc.)

    • Edge or on-device inference experience (llama.cpp, MLX, ONNX Runtime, etc.)

    • Experience with quantization, speculative decoding, disaggregated inference or KV-cache compression

    • Experience with multimodal models and/or agentic systems

Location:
  • Paris or London.

  • This role is hybrid, and you are expected to be in the office 3 days a week on average.

  • Please expect some travel between offices on a reasonable cadence (e.g., every 4-6 weeks).

What We Offer:
  • Join the exciting journey of shaping the future of AI

  • Collaborate with a fun, dynamic and multicultural team, working alongside world-class AI talent in a highly collaborative environment

  • Enjoy a competitive salary

  • Unlock opportunities for professional growth, continuous learning, and career development

If you want to change the status quo in AI, join us.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Research Engineer, Model Inference & Serving - Paris
Research Engineer, Model Inference & Serving - Paris

H Company • Paris

Sur place
EUR 60 000 - 90 000
Competitive salary
Opportunities for professional growth
Collaborative multicultural team
Research Engineer / Scientist, Post-training - Paris
Research Engineer / Scientist, Post-training - Paris

hcompany • Paris

Hybride
EUR 90 000 - 130 000
Hybrid work model
Travel between offices
Professional growth opportunities
+1
Member of technical staff (Infrastructure) - Paris
Member of technical staff (Infrastructure) - Paris

hcompany • Paris

Hybride
EUR 90 000 - 120 000
Competitive salary
Career development
Collaborative team
Research Engineer / Scientist, Post-training - Paris
Research Engineer / Scientist, Post-training - Paris

H Company • Paris

Hybride
EUR 70 000 - 100 000
Competitive salary
Opportunities for professional growth
Collaborative and dynamic work environment
Member of technical staff (Infrastructure) - Paris
Member of technical staff (Infrastructure) - Paris

H Company • Paris

Hybride
EUR 90 000 - 130 000
Competitive salary
Collaborative team environment
Career growth opportunities
Member of technical staff (Infrastructure) - Paris
Member of technical staff (Infrastructure) - Paris

H Company • Paris

Hybride
EUR 90 000 - 130 000
Competitive salary
Career growth and professional develop
Member of technical staff (Infrastructure) - London
Member of technical staff (Infrastructure) - London

H Company • Paris

Hybride
EUR 90 000 - 120 000
In-office collaboration
Forward Deployed Engineer
Forward Deployed Engineer

H Company • Paris

Hybride
EUR 90 000 - 130 000
Senior Talent Acquisition partner
Senior Talent Acquisition partner

H Company • Paris

Hybride
EUR 90 000 - 130 000
Health insurance (100% coverage)
Equity
Hybrid work model
+2
ML Infrastructure Engineer
ML Infrastructure Engineer

Npv • Paris

Hybride
EUR 120 000 - 190 000
Competitive salary + equity
Hybrid work from Paris with relocation
Top-tier medical insurance in France
+2