Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Xist4 IT Limited. in the United Kingdom is hiring an ML Platform Engineer to own platform infrastructure for training, evaluation, deployment and observation. Expect reusable, scalable systems and close collaboration with AI engineers and researchers.
You will optimize inference pipelines for high throughput and low latency, work with vLLM, SGLang and TensorRT-LLM, and drive observability across GPU and distributed systems to prevent regressions.
£90,000 to £115,000. Permanent.
London. Fully remote across the UK.
The models are only useful if the team can train, evaluate, deploy and observe them without rebuilding the machinery each time. You will own that machinery.
Our client is an early-stage AI product company building applications that get on with everyday tasks before you ask. A prototype exists, launch is ahead, and the company is funded without relying on an upcoming round.
You'll work across the infrastructure behind the AI stack, alongside AI engineers, researchers and product engineers. The aim is to make experimentation quick, production releases dependable and the expensive parts of the stack visible.
This is infrastructure work with ML consequences. Throughput, latency, GPU use, cost, model regressions and failed releases all land here. You need to be comfortable debugging across the line between platform and model serving.
Platform. Build and operate the systems used for model training, evaluation, deployment, inference and experimentation. The useful outcome is reusable infrastructure that other engineers can rely on, not a pile of one-off pipelines.
Serving. Build and optimise inference infrastructure for high-throughput, low-latency workloads. You'll work with serving technologies such as vLLM, SGLang or TensorRT-LLM and find the bottlenecks across GPU and distributed systems.
Pipelines. Develop repeatable flows for data preparation, training, evaluation, model release and continuous improvement. Reproducibility and maintainability matter as much as getting a first run through.
Observability. Build monitoring, tracing, alerting and benchmarking for AI workloads. You'll make model and infrastructure regressions visible early enough to do something about them.
PyTorch, JAX, vLLM, SGLang, TensorRT-LLM, GPU tooling, vector databases, workflow orchestration.
You are probably an ML Platform Engineer, MLOps Engineer or Machine Learning Infrastructure Engineer who likes building the common layer underneath model teams. You think about the whole route from experiment to serving.
It won't suit you if you want to run infrastructure and leave the models to someone else. Here a slow endpoint or a regression after a release is yours to trace, through the GPUs, the serving layer and the model itself.
We welcome applicants from every background and will support reasonable adjustments.