ML Engineer: Real-Time Distributed Training & Inference

IMC

City Of London

On-site

GBP 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

IMC is a research-driven trading firm leveraging machine learning to advance its trading capabilities. We seek a Machine Learning Engineer to build large-scale training and real-time inference pipelines, collaborate with researchers and HPC specialists, and push the performance of GPU-accelerated systems.

You will work on distributed training, low-latency inference, and open-source ML tool integration to drive faster experimentation and real-time decision making in trading environments.

Qualifications

  • 5+ years in ML with focus on training or inference systems.
  • Real-time, low-latency ML pipelines in high-performance environments a strong plus.
  • Proficiency in Python, CUDA, or C++.
  • Knowledge of PyTorch, TensorFlow or JAX.
  • GPU programming for training and inference acceleration (CuDNN, TensorRT).
  • Experience with distributed training (Horovod, NCCL).
  • Exposure to cloud platforms and orchestration tools.
  • Open-source contributions in ML or distributed systems a plus.

Responsibilities

  • Develop large-scale distributed training pipelines for datasets and models.
  • Build and optimize low-latency inference pipelines for real-time predictions.
  • Develop libraries to improve ML framework performance.
  • Maximize training and inference performance using GPUs and acceleration libraries.
  • Design scalable model frameworks for high-volume trading data and real-time predictions.
  • Collaborate with researchers to automate ML experiments and model retraining.
  • Partner with HPC specialists to optimize workflows and reduce costs.
  • Evaluate and roll out third-party tools to enhance development and inference.
  • Explore open-source ML tool internals to extend capabilities.

Skills

Python
CUDA
C++
PyTorch
TensorFlow
JAX
Horovod
NCCL
GPU acceleration
Distributed systems

Tools

CuDNN
TensorRT
NCCL
Horovod
CUDA toolkit
MPI

Job description

IMC is a research-driven trading firm leveraging machine learning to advance its trading capabilities. We seek a Machine Learning Engineer to build large-scale training and real-time inference pipelines, collaborate with researchers and HPC specialists, and push the performance of GPU-accelerated systems.

You will work on distributed training, low-latency inference, and open-source ML tool integration to drive faster experimentation and real-time decision making in trading environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Engineer: Distributed Training & Low-Latency Inference
ML Engineer: Distributed Training & Low-Latency Inference

IMC Trading • Greater London

On-site
GBP 70,000 - 90,000
Machine Learning Engineer
Machine Learning Engineer

IMC • City Of London

On-site
GBP 120,000 - 160,000
Machine Learning Engineer
Machine Learning Engineer

IMC Trading • Greater London

On-site
GBP 70,000 - 90,000
ML Research Engineer - Finance, HPC & Low-Latency
ML Research Engineer - Finance, HPC & Low-Latency

Tradermath • Greater London

Hybrid
GBP 90,000 - 130,000
Private Medical Insurance
Vision and Dental Insurance
Group Pension Scheme
+3
Graduate ML Researcher — Trading Analytics
Graduate ML Researcher — Trading Analytics

targetjobs UK • Greater London

On-site
GBP 45,000 - 65,000
Senior Real-Time ML Inference Engineer
Senior Real-Time ML Inference Engineer

BITKRAFT Ventures • United Kingdom

On-site
GBP 140,000 - 200,000
ML Infrastructure Engineer - Scale AI Training & Inference
ML Infrastructure Engineer - Scale AI Training & Inference

Isomorphic Labs • City of Westminster

On-site
GBP 90,000 - 130,000
Campus ML Engineer: Build Scalable AI for Finance
Campus ML Engineer: Build Scalable AI for Finance

Trading Interview • Greater London

Hybrid
GBP 90,000 - 140,000
ML Research Engineer – Financial ML & Low-Latency GPU
ML Research Engineer – Financial ML & Low-Latency GPU

Amsterdam Quant Society • Greater London

Hybrid
GBP 70,000 - 110,000
Principal Research Scientist – Machine Learning
Principal Research Scientist – Machine Learning

Trading Interview • Greater London

On-site
GBP 120,000 - 180,000
Competitive compensation
World‑class datasets and compute
Conference and publication support