Multimodal AI Inference & Serving Engineer

H

Paris (TX)

Hybrid

USD 104,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary
Career development
Multicultural team

Job summary

H invites a Research Engineer to advance the inference and serving stack for its multimodal, agentic AI models. This hybrid role is based in Paris or London, with in-office presence three days a week and occasional travel.

You will optimize latency, throughput and cost, prototype inference techniques, and collaborate with the Models team on training-time decisions that affect deployment. A strong software background, ML experience, and the ability to communicate across teams are essential.

Qualifications

  • Strong software engineering background with production experience.
  • Proficient in Python and at least one systems language (Rust, C++, or Go).
  • Hands-on experience with deep learning frameworks (PyTorch, JAX) in industry.
  • Solid distributed systems fundamentals.
  • Experience with modern cloud environments and ML infra (Kubernetes, etc.).
  • Knowledge of transformers and multimodal architectures.

Responsibilities

  • Build and operate the inference stack for multimodal agentic models.
  • Improve latency, throughput, and cost of model serving.
  • Research inference techniques for agent workloads.
  • Collaborate with Models team on training-time decisions.
  • Integrate inference into AI products with cross-functional teams.
  • Evaluate platforms and communicate findings to stakeholders.
  • Stay current with advancements in inference and accelerator tech.

Skills

Strong software engineering
Python
Rust
C++
Go
Deep learning frameworks
Distributed systems
Cloud infra
ML model serving

Education

Advanced degree (PhD/MS) with research output

Tools

vLLM
SGLang
TensorRT-LLM
CUDA

Job description

H invites a Research Engineer to advance the inference and serving stack for its multimodal, agentic AI models. This hybrid role is based in Paris or London, with in-office presence three days a week and occasional travel.

You will optimize latency, throughput and cost, prototype inference techniques, and collaborate with the Models team on training-time decisions that affect deployment. A strong software background, ML experience, and the ability to communicate across teams are essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Model Inference & Serving - Paris
Research Engineer, Model Inference & Serving - Paris

H • Paris (TX)

Hybrid
USD 104,000 - 150,000
Competitive salary
Career development
Multicultural team
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Boston (MA), Northern (KY)

Hybrid
USD 167,000 - 226,000
Senior Real-Time Multimodal Inference Architect
Senior Real-Time Multimodal Inference Architect

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
Staff AI Inference & Systems Engineer
Staff AI Inference & Systems Engineer

Mixpeek • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Multimodal Inference Engineer — Scale GPU AI Models
Multimodal Inference Engineer — Scale GPU AI Models

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Researcher, Multimodal AI Architect
Researcher, Multimodal AI Architect

Cartesia • United States

Hybrid
USD 159,000 - 225,000
Compensation
Commuter allowance
Flexible PTO
+2
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+2
Multimodal AI Research Scientist
Multimodal AI Research Scientist

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3