Research Engineer, Model Inference & Serving - Paris

H

Paris (TX)

Hybrid

USD 104,000 - 150,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive salary
Career development
Multicultural team

Job summary

H invites a Research Engineer to advance the inference and serving stack for its multimodal, agentic AI models. This hybrid role is based in Paris or London, with in-office presence three days a week and occasional travel.

You will optimize latency, throughput and cost, prototype inference techniques, and collaborate with the Models team on training-time decisions that affect deployment. A strong software background, ML experience, and the ability to communicate across teams are essential.

Qualifications

  • Strong software engineering background with production experience.
  • Proficient in Python and at least one systems language (Rust, C++, or Go).
  • Hands-on experience with deep learning frameworks (PyTorch, JAX) in industry.
  • Solid distributed systems fundamentals.
  • Experience with modern cloud environments and ML infra (Kubernetes, etc.).
  • Knowledge of transformers and multimodal architectures.

Responsibilities

  • Build and operate the inference stack for multimodal agentic models.
  • Improve latency, throughput, and cost of model serving.
  • Research inference techniques for agent workloads.
  • Collaborate with Models team on training-time decisions.
  • Integrate inference into AI products with cross-functional teams.
  • Evaluate platforms and communicate findings to stakeholders.
  • Stay current with advancements in inference and accelerator tech.

Skills

Strong software engineering
Python
Rust
C++
Go
Deep learning frameworks
Distributed systems
Cloud infra
ML model serving

Education

Advanced degree (PhD/MS) with research output

Tools

vLLM
SGLang
TensorRT-LLM
CUDA

Job description

Research Engineer, Model Inference & Serving

About H: H exists to push the boundaries of superintelligence with agentic AI. By automating complex, multi-step tasks typically performed by humans, AI agents will help unlock full human potential. H is hiring the world's best AI talent, seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute.

About the Team: The Inference team builds and operates the systems that serve H's foundational models in production. We focus on multimodal inference and serving for Computer Use Agents, optimizing across both the inference engine layer (e.g., vLLM, SGLang) and the model serving layer (e.g., disaggregated inference, intelligent routing). Agentic inference brings constraints around context length, multimodality, and tool calls, which we address by co-designing with the Models team on training-time choices and with the agent teams on how models are deployed. We operate at the intersection of research and production, translating cutting-edge inference techniques into the systems that power H's next generation of agents. We are looking for strong engineers excited about inference to join the team and help shape the systems behind superintelligent AI.

Key Responsibilities:
  • Build and operate the inference stack that serves H's multimodal agentic models

  • Improve latency, throughput, and cost of model serving across the stack

  • Research and implement inference techniques tailored to agent workloads

  • Co-design with the Models team on training-time decisions that affect inference

  • Collaborate with cross-functional teams to integrate inference into agentic AI products

  • Evaluate inference, serving, and hardware platforms, and communicate findings to stakeholders

  • Stay current with advancements in inference, model serving, and accelerator technology

Requirements:
  • Technical skills:

    • Strong software engineering track record

    • Proficient in Python and at least one systems language (Rust, C++, or Go)

    • Hands‑on experience with deep learning frameworks (PyTorch, JAX), preferably in an industry setting

    • Solid distributed systems fundamentals

    • Experience working in a modern cloud environment and with production ML infrastructure (Kubernetes, etc.)

    • Working knowledge of modern ML, including transformers and multimodal architectures

  • Research skills:

    • Research engagement: an advanced degree with research output, or publications at top-tier AI or systems venues (e.g., NeurIPS, ICML, MLSys, OSDI), research internships, or substantive open-source contributions

  • Soft skills:

    • Excellent communication and presentation skills

    • Strong collaboration and teamwork skills

    • Passion for inference and AI

  • Preferred qualifications:

    • Startup experience

    • Hands‑on experience with inference frameworks (vLLM, SGLang, TensorRT-LLM)

    • Writing or modifying GPU kernels (CUDA, Triton, etc.)

    • Edge or on-device inference experience (llama.cpp, MLX, ONNX Runtime, etc.)

    • Experience with quantization, speculative decoding, disaggregated inference or KV-cache compression

    • Experience with multimodal models and/or agentic systems

Location:
  • Paris or London.

  • This role is hybrid, and you are expected to be in the office 3 days a week on average.

  • Please expect some travel between offices on a reasonable cadence (e.g., every 4-6 weeks).

What We Offer:
  • Join the exciting journey of shaping the future of AI

  • Collaborate with a fun, dynamic and multicultural team, working alongside world‑class AI talent in a highly collaborative environment

  • Enjoy a competitive salary

  • Unlock opportunities for professional growth, continuous learning, and career development

If you want to change the status quo in AI, join us.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Staff + Senior Software Engineer, Inference Deployment
Staff + Senior Software Engineer, Inference Deployment

Anthropic • United States

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation
Parental leave
+2
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Staff Software Engineer - Inference Backends
Staff Software Engineer - Inference Backends

Hume AI • New York (NY)

On-site
USD 140,000 - 190,000
Multimodal AI Inference & Serving Engineer
Multimodal AI Inference & Serving Engineer

H • Paris (TX)

Hybrid
USD 104,000 - 150,000
Competitive salary
Career development
Multicultural team
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Engineering Manager, Inference Infrastructure
Engineering Manager, Inference Infrastructure

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 625,000
Competitive compensation
Equity matching
Parental leave
+2
Engineering Manager, Inference Infrastructure
Engineering Manager, Inference Infrastructure

Anthropic • New York (NY)

Hybrid
USD 405,000 - 625,000
Competitive compensation
Equity donation
Vacation and parental leave
+2
Performance Engineer, Inference Engine
Performance Engineer, Inference Engine

Anthropic • San Francisco (CA), New York (NY)

On-site
USD 350,000 - 850,000
Forward Deployed Engineer
Forward Deployed Engineer

H Company • United States

Hybrid
USD 150,000 - 200,000
Equity