Research Software Engineer: Scale Distributed ML Pipelines
Deepstreamtech
Greater London
On-site
GBP 50,000 - 80,000
Full time
14 days+
Application generator
An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Get past ATS filters
Job summary
Deepstreamtech is looking for a Research Engineer based in London to design and develop distributed systems for training frontier-scale models. You'll be responsible for writing clean, reliable code in Python and systems languages, while maintaining tools and internal pipelines. The ideal candidate has a Master’s in Computer Science and over 4 years of experience in large-scale systems development. Exposure to machine learning workflows is a plus. Join a collaborative team committed to innovation.
Qualifications
4+ years building and operating large-scale or distributed systems.
Hands-on with container orchestration and schedulers (Kubernetes / K8s, SLURM, or similar).
Responsibilities
Design and harden the codebase, tools, and distributed services for training frontier-scale models.
Build and maintain shared dev-tools, evaluation & data pipelines, training framework, and cluster tooling.
Interface research with product: expose clean APIs, automate model pushes, surface live metrics.
Skills
Python
C++
Rust
Go
Java
Kubernetes
SLURM
Education
Master’s in Computer Science or equivalent experience
Job description
Deepstreamtech is looking for a Research Engineer based in London to design and develop distributed systems for training frontier-scale models. You'll be responsible for writing clean, reliable code in Python and systems languages, while maintaining tools and internal pipelines. The ideal candidate has a Master’s in Computer Science and over 4 years of experience in large-scale systems development. Exposure to machine learning workflows is a plus. Join a collaborative team committed to innovation.