Inference Platform Engineer — Scalable LLM Serving

Mistral

Greater London

Híbrido

GBP 120 000 - 160 000

Tempo integral

há 17 horas
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Destaca-te nesta função — gera um currículo e uma carta de apresentação personalizados em cerca de um minuto.

Ultrapassa os filtros ATS

Resumo da oferta

Mistral is hiring for an Inference Foundation engineer to own the core of the inference stack, from the engine to orchestration, and to deliver production-grade serving for frontier models.

This hybrid role covers production LLM serving, engine and platform development, and capacity engineering. You will tackle scalable optimization, elastic capacity, and high-throughput training pipelines while collaborating across distributed teams.

Qualificações

  • Experience building and running ML/LLM services at scale with latency targets.
  • Hands-on experience with inference engines such as vLLM, sgLang, TensorRT-LLM.
  • Familiarity with distributed serving architectures and GPU networking.

Responsabilidades

  • Develop and fix the core of the inference stack — engine and orchestrator — including feature selection, configuration, and tuning for maximum performance at scale
  • Own the release process for the serving stack: validated, regression-free releases through automated performance gates and progressive rollout
  • Drive improvements and fixes upstream when the open-source engine is the right place for them
  • Drive optimization of serving efficiency and topology at scale—reducing startup times, cache misses, and data movement

Conhecimentos

ML/LLM services at scale
Inference engines
Distributed serving architectures
CUDA/NCCL debugging
Python backend tooling
Kubernetes
GPU networking basics
Rust or C++ production

Ferramentas

PyTorch
CUDA
NCCL
Docker
Kubernetes
Nsight Systems/Compute
Profiling tools

Descrição da oferta de emprego

Mistral is hiring for an Inference Foundation engineer to own the core of the inference stack, from the engine to orchestration, and to deliver production-grade serving for frontier models.

This hybrid role covers production LLM serving, engine and platform development, and capacity engineering. You will tackle scalable optimization, elastic capacity, and high-throughput training pipelines while collaborating across distributed teams.

Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

AI Infra Engineer: Scalable LLM Serving & Platform Design
AI Infra Engineer: Scalable LLM Serving & Platform Design

scaleai • Greater London

Presencial
GBP 110 000 - 160 000
ML Platform Lead - LLM Training & Inference
ML Platform Lead - LLM Training & Inference

Scale AI • York and North Yorkshire

Presencial
GBP 120 000 - 180 000
Health & Wellbeing
Career Growth stipend
Community events
+1
Senior LLM Serving Platform Engineer
Senior LLM Serving Platform Engineer

Scale AI • Greater London

Presencial
GBP 90 000 - 130 000
Research Engineer, Inference Foundation
Research Engineer, Inference Foundation

Mistral • Greater London

Híbrido
GBP 120 000 - 160 000
Senior Data Infrastructure Engineer, AI Compute Platform
Senior Data Infrastructure Engineer, AI Compute Platform

Mistral • Greater London

Híbrido
GBP 76 673 - 110 750
Healthcare coverage
Relocation support
Wellness programs
Senior AI Platform Engineer: Scalable LLM Infrastructure
Senior AI Platform Engineer: Scalable LLM Infrastructure

IQVIA LLC • Greater London

Presencial
GBP 120 000 - 180 000
Research Engineer - Scalable ML & Production Tools
Research Engineer - Scalable ML & Production Tools

Mistral • Greater London

Presencial
GBP 100 000 - 140 000
Healthcare coverage
Parental leave
Retirement plans
+3
AI Infrastructure Engineer, Serving Platform
AI Infrastructure Engineer, Serving Platform

Scale AI • Greater London

Presencial
GBP 90 000 - 130 000
Staff Engineer, Inference Platform for AI Drug Discovery
Staff Engineer, Inference Platform for AI Drug Discovery

Isomorphic Labs • Greater London

Híbrido
GBP 120 000 - 150 000
Hybrid work model
Platform Backend Engineer: Scalable Infra & LLM Systems
Platform Backend Engineer: Scalable Infra & LLM Systems

Conduct AI • Greater London

Híbrido
GBP 90 000 - 130 000