Research Engineer, Model Inference & Serving - London

hcompany

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary
Hybrid work model
Career development

Job summary

H is seeking a Research Engineer to advance inference and serving for multimodal agentic AI. You will build and operate the production stack, optimizing latency and cost while collaborating with the Models team and cross-functional partners in a hybrid Paris/London setting.

You will apply Python and systems programming, work with PyTorch/JAX, and contribute to scalable ML infrastructure across cloud environments. This role offers growth, a competitive salary, and travel between offices.

Qualifications

  • Strong software engineering track record.
  • Proficient in Python and at least one systems language such as Rust, C++, or Go.
  • Hands-on experience with deep learning frameworks (PyTorch, JAX).
  • Solid distributed systems fundamentals.
  • Experience with production ML infrastructure (Kubernetes).
  • Working knowledge of modern ML including transformers and multimodal architectures.
  • Advanced degree with research output or publications in AI/systems venues.

Responsibilities

  • Build and operate the inference stack that serves H's multimodal agentic models.
  • Improve latency, throughput, and cost of model serving across the stack.
  • Research and implement inference techniques tailored to agent workloads.
  • Co-design with the Models team on training-time decisions that affect inference.
  • Collaborate with cross-functional teams to integrate inference into agentic AI products.
  • Evaluate inference, serving, and hardware platforms, and communicate findings to stakeholders.
  • Stay current with advancements in inference, model serving, and accelerator technology.

Skills

Python
Rust
C++
Go
PyTorch
JAX
Distributed systems
Kubernetes
Cloud infra
Transformers
Multimodal models

Education

Advanced degree

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
Triton
llama.cpp
MLX
ONNX Runtime

Job description

Research Engineer, Model Inference & Serving

About H: H exists to push the boundaries of superintelligence with agentic AI. By automating complex, multi-step tasks typically performed by humans, AI agents will help unlock full human potential. H is hiring the world's best AI talent, seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute.

About the Team: The Inference team builds and operates the systems that serve H's foundational models in production. We focus on multimodal inference and serving for Computer Use Agents, optimizing across both the inference engine layer (e.g., vLLM, SGLang) and the model serving layer (e.g., disaggregated inference, intelligent routing). Agentic inference brings constraints around context length, multimodality, and tool calls, which we address by co-designing with the Models team on training-time choices and with the agent teams on how models are deployed. We operate at the intersection of research and production, translating cutting-edge inference techniques into the systems that power H's next generation of agents. We are looking for strong engineers excited about inference to join the team and help shape the systems behind superintelligent AI.

Key Responsibilities:
  • Build and operate the inference stack that serves H's multimodal agentic models

  • Improve latency, throughput, and cost of model serving across the stack

  • Research and implement inference techniques tailored to agent workloads

  • Co-design with the Models team on training-time decisions that affect inference

  • Collaborate with cross-functional teams to integrate inference into agentic AI products

  • Evaluate inference, serving, and hardware platforms, and communicate findings to stakeholders

  • Stay current with advancements in inference, model serving, and accelerator technology

Requirements:
  • Technical skills:

    • Strong software engineering track record

    • Proficient in Python and at least one systems language (Rust, C++, or Go)

    • Hands-on experience with deep learning frameworks (PyTorch, JAX), preferably in an industry setting

    • Solid distributed systems fundamentals

    • Experience working in a modern cloud environment and with production ML infrastructure (Kubernetes, etc.)

    • Working knowledge of modern ML, including transformers and multimodal architectures

  • Research skills:

    • Research engagement: an advanced degree with research output, or publications at top-tier AI or systems venues (e.g., NeurIPS, ICML, MLSys, OSDI), research internships, or substantive open-source contributions

  • Soft skills:

    • Excellent communication and presentation skills

    • Strong collaboration and teamwork skills

    • Passion for inference and AI

  • Preferred qualifications:

    • Startup experience

    • Hands-on experience with inference frameworks (vLLM, SGLang, TensorRT-LLM)

    • Writing or modifying GPU kernels (CUDA, Triton, etc.)

    • Edge or on-device inference experience (llama.cpp, MLX, ONNX Runtime, etc.)

    • Experience with quantization, speculative decoding, disaggregated inference or KV-cache compression

    • Experience with multimodal models and/or agentic systems

Location:
  • Paris or London.

  • This role is hybrid, and you are expected to be in the office 3 days a week on average.

  • Please expect some travel between offices on a reasonable cadence (e.g., every 4-6 weeks).

What We Offer:
  • Join the exciting journey of shaping the future of AI

  • Collaborate with a fun, dynamic and multicultural team, working alongside world-class AI talent in a highly collaborative environment

  • Enjoy a competitive salary

  • Unlock opportunities for professional growth, continuous learning, and career development

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer / Scientist, Post-training - London
Research Engineer / Scientist, Post-training - London

H Company • Greater London

Hybrid
GBP 60,000 - 90,000
Competitive salary
Opportunities for professional growth
Collaborative and multicultural team environment
Research Engineer / Scientist, Post-training - London
Research Engineer / Scientist, Post-training - London

H Company • Greater London

Hybrid
GBP 120,000 - 180,000
Research Engineer / Scientist, Post-training - Paris
Research Engineer / Scientist, Post-training - Paris

Creandum • Greater London

Hybrid
GBP 110,000 - 150,000
Member of technical staff (Infrastructure) - London
Member of technical staff (Infrastructure) - London

hcompany • Greater London

Hybrid
GBP 70,000 - 110,000
Hybrid work model
Competitive salary
Member of technical staff (Infrastructure) - Paris
Member of technical staff (Infrastructure) - Paris

Creandum • Greater London

Hybrid
GBP 90,000 - 120,000
Competitive salary
Career growth opportunities
Dynamic multicultural team
Staff Software Engineer, Inference
Staff Software Engineer, Inference

Mat Vin • Greater London

Hybrid
GBP 325,000 - 390,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Research Engineer (Inference & Serving)
Research Engineer (Inference & Serving)

Axiōma Search • Greater London

On-site
GBP 90,000 - 120,000
Research Engineer, Multimodal Inference & Serving
Research Engineer, Multimodal Inference & Serving

hcompany • Greater London

Hybrid
GBP 90,000 - 130,000
Competitive salary
Hybrid work model
Career development
AI Engineer
AI Engineer

G-Research • Greater London

On-site
GBP 90,000 - 150,000
Highly competitive compensation
Annual discretionary bonus
Lunch provided
+5
Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)
Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Good lunch and dinner
Collaborative work culture
No bureaucracy