ML Platform Engineer

Xist4 IT Limited.

City Of London

Remote

GBP 90,000 - 115,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Xist4 IT Limited. in the United Kingdom is hiring an ML Platform Engineer to own platform infrastructure for training, evaluation, deployment and observation. Expect reusable, scalable systems and close collaboration with AI engineers and researchers.

You will optimize inference pipelines for high throughput and low latency, work with vLLM, SGLang and TensorRT-LLM, and drive observability across GPU and distributed systems to prevent regressions.

Qualifications

  • Experience building/operating ML infrastructure or production ML platforms.
  • Production Python used for platform or ML systems.
  • Ownership of deployment, inference or evaluation infrastructure.

Responsibilities

  • Platform: Build and operate systems for training, evaluation, deployment, inference and experimentation.
  • Serving: Optimise inference infrastructure for high-throughput, low-latency workloads with vLLM/SGLang/TensorRT-LLM.
  • Pipelines: Develop repeatable data prep, training, evaluation, model release and continuous improvement flows.
  • Observability: Build monitoring, tracing, alerting and benchmarking for ML workloads.

Skills

ML infrastructure
Production Python
Model deployment
Distributed systems
Low-latency serving
Data pipelines
Production monitoring
UK work eligibility

Tools

PyTorch
JAX
vLLM
SGLang
TensorRT-LLM
GPU tooling
Vector databases
Workflow orchestration

Job description

ML Platform Engineer | Python, model serving, distributed systems | Fully remote, UK

£90,000 to £115,000. Permanent.
London. Fully remote across the UK.

The models are only useful if the team can train, evaluate, deploy and observe them without rebuilding the machinery each time. You will own that machinery.

Our client is an early-stage AI product company building applications that get on with everyday tasks before you ask. A prototype exists, launch is ahead, and the company is funded without relying on an upcoming round.

You'll work across the infrastructure behind the AI stack, alongside AI engineers, researchers and product engineers. The aim is to make experimentation quick, production releases dependable and the expensive parts of the stack visible.

This is infrastructure work with ML consequences. Throughput, latency, GPU use, cost, model regressions and failed releases all land here. You need to be comfortable debugging across the line between platform and model serving.

The job

Platform. Build and operate the systems used for model training, evaluation, deployment, inference and experimentation. The useful outcome is reusable infrastructure that other engineers can rely on, not a pile of one-off pipelines.
Serving. Build and optimise inference infrastructure for high-throughput, low-latency workloads. You'll work with serving technologies such as vLLM, SGLang or TensorRT-LLM and find the bottlenecks across GPU and distributed systems.
Pipelines. Develop repeatable flows for data preparation, training, evaluation, model release and continuous improvement. Reproducibility and maintainability matter as much as getting a first run through.
Observability. Build monitoring, tracing, alerting and benchmarking for AI workloads. You'll make model and infrastructure regressions visible early enough to do something about them.

What you'll bring
  • ML infrastructure or production ML platforms you have built or operated.
  • Production Python used for platform or machine learning systems.
  • Model deployment, inference or evaluation infrastructure you have owned.
  • Distributed systems work where reliability mattered in production.
  • Low-latency or high-throughput serving you have tuned.
  • ML or data pipelines built to be repeatable and observable.
  • Production monitoring, tracing or alerting for ML workloads.
  • Existing right to work in the UK.
Useful:

PyTorch, JAX, vLLM, SGLang, TensorRT-LLM, GPU tooling, vector databases, workflow orchestration.

Who this suits

You are probably an ML Platform Engineer, MLOps Engineer or Machine Learning Infrastructure Engineer who likes building the common layer underneath model teams. You think about the whole route from experiment to serving.

It won't suit you if you want to run infrastructure and leave the models to someone else. Here a slow endpoint or a regression after a release is yours to trace, through the GPUs, the serving layer and the model itself.

Let’s talk

We welcome applicants from every background and will support reasonable adjustments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Xist4 IT Limited. • City Of London

Remote
GBP 95,000 - 120,000
Machine Learning Engineer
Machine Learning Engineer

Xist4 IT Limited. • City Of London

Remote
GBP 65,000 - 85,000
Platform Engineer
Platform Engineer

Stealth iT Consulting • Greater London, Manchester, Glasgow

Hybrid
GBP 42,000 - 70,000
Bonus & Benefits
Engineering manager - ML Platform
Engineering manager - ML Platform

Velocity Tech • England

On-site
GBP 85,000 - 110,000
Work with the latest AI technology
Help build products used by millions
Friendly team environment
+1
Senior MLOps Engineer - AI Infrastructure
Senior MLOps Engineer - AI Infrastructure

Harnham • Greater London

Hybrid
GBP 51,000 - 85,000
Hybrid work model
ML Systems Engineer (RL)
ML Systems Engineer (RL)

Roc Search Inc. • Greater London

Hybrid
GBP 100,000 - 110,000
Senior ML Ops Engineer
Senior ML Ops Engineer

Harnham - Data and Analytics Recruitment • Greater London

Hybrid
GBP 75,000 - 85,000
Hybrid work model
Machine Learning Manager
Machine Learning Manager

SPG Resourcing • York and North Yorkshire

On-site
GBP 70,000 - 90,000
Senior ML Ops Engineer 201043
Senior ML Ops Engineer 201043

Harnham • United Kingdom

On-site
GBP 70,000 - 110,000
Ongoing training
Ownership of ML Ops function
DevOps and Machine Learning Operations Engineer
DevOps and Machine Learning Operations Engineer

Kintec Global Recruitment • Manchester

On-site
GBP 70,000 - 90,000