Member of Technical Staff

Salient Group

Greater London

Hybrid

GBP 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity grant

Job summary

Salient Group is hiring for a Member of Technical Staff focusing on AI infrastructure / ML platform in London. The role involves building a research platform that connects models, compute and data with scientists, enabling fast, reproducible experiments.

You will own GPU infrastructure for training and inference, design scalable data pipelines, and productionise research workflows. Hybrid London working is offered with meaningful equity in a well-funded environment.

Qualifications

  • Proficient Python and strong software engineering fundamentals.
  • Experience building ML infrastructure, ML platforms, MLOps or distributed systems in production.
  • Understanding of modern ML lifecycle: data, training, evaluation, deployment, inference and monitoring.
  • Experience with GPU workloads and performance/reliability challenges.
  • Experience with containerisation and orchestration (Docker, Kubernetes).
  • Strong knowledge of cloud infrastructure and Infrastructure-as-Code.
  • Experience designing reliable data pipelines, APIs and distributed services.
  • Ability to reason from first principles about requirements, scale and trade-offs.

Responsibilities

  • Build the research platform connecting models, compute, data and scientific workflows for reproducible experiments.
  • Own AI infrastructure: GPU training and inference with emphasis on scheduling, latency, throughput and cost.
  • Develop scalable data pipelines for ingesting and serving scientific data.
  • Productionise research: move ideas from notebooks to robust systems without excessive overhead.
  • Improve model serving by profiling and optimizing inference workloads.

Skills

Python
ML infrastructure
GPU workloads
Docker
Kubernetes
Cloud IAC
Distributed systems

Tools

PyTorch
TensorRT

Job description

Member of Technical Staff - AI Infrastructure / ML Platform

Location | London (Hybrid)

Comp | Highly Competitive + Meaningful Equity

Focus | AI Infrastructure, MLOps, Data Platform, GPU Systems, Research Infrastructure, Scientific ML

We’re partnering with a stealth AI4Science company building foundational technology designed to accelerate science.

Backed by significant funding and leading technology investors, the company is bringing together exceptional researchers, engineers and scientists to tackle problems where advances in AI can translate into discoveries in the physical world.

Unlike traditional software environments, the infrastructure here sits directly underneath a scientific research engine: large-scale model training and inference, simulation, experimental data, GPU workloads and research workflows all need to work together reliably.

They are now looking for a Member of Technical Staff focused on AI Infrastructure / ML Platform to help build that foundation.

This is an early and highly influential hire. You’ll work directly with researchers and scientists, understand how they actually experiment, and build the systems that allow them to move significantly faster.

About The Role

As a Member of Technical Staff, you’ll own infrastructure across the intersection of ML systems, data, compute and research engineering.

The challenge isn’t to build an enormous internal platform for its own sake. It’s to understand what researchers need, identify the bottlenecks slowing them down, and build the smallest, strongest abstractions that make experimentation faster and more reliable.

You could be working on GPU orchestration one week, research data infrastructure the next, and improving model serving, experiment reproducibility or distributed training workflows after that.

You’ll have significant freedom to make build‑vs‑buy decisions, introduce new infrastructure where it creates genuine leverage, and deliberately avoid unnecessary complexity where it doesn’t.

Examples Of The Problems You Might Tackle
  • Building the ML platform researchers use to train, evaluate, deploy and iterate on scientific models.
  • Designing infrastructure for GPU‑intensive training and inference workloads, including scheduling, resource utilisation and workload isolation.
  • Improving distributed model serving using technologies such as vLLM, TensorRT, Triton, Ray or equivalent systems.
  • Building reliable data pipelines and storage systems connecting simulation, modelling, experimental data and downstream research workflows.
  • Creating reproducible experimentation environments so researchers can move quickly without repeatedly solving infrastructure problems.
  • Designing systems for experiment tracking, model versioning, evaluation, observability and lineage.
  • Improving inference latency, throughput, GPU utilisation and cost efficiency.
  • Building internal APIs, tooling and abstractions that make complex infrastructure accessible to researchers without constraining how they work.
  • Supporting workloads that may move between local compute, cloud infrastructure and dedicated GPU environments.
  • Designing systems capable of evolving as the organisation moves from individual research experiments towards increasingly automated scientific workflows.
What You’ll Do
  • Build the research platform: Design the infrastructure connecting models, compute, data and scientific workflows, allowing researchers to move from an idea to a reproducible experiment quickly.
  • Own AI infrastructure: Build and operate GPU infrastructure for model training and inference, thinking carefully about scheduling, utilisation, latency, throughput, reliability and cost.
  • Build the data layer: Develop scalable pipelines and systems for ingesting, processing, storing and serving scientific, simulation and experimental data.
  • Productionise research: Help researchers move promising ideas beyond notebooks into robust systems without introducing unnecessary process or infrastructure overhead.
  • Improve model serving: Profile and optimise inference workloads, choosing the right serving architecture and hardware configuration for different models and research requirements.
  • Design for researchers: Work directly with Research Scientists and Engineers to understand how they work and build tools that increase their velocity rather than forcing them into rigid platform abstractions.
  • Make pragmatic architecture decisions: Start from requirements, scale and constraints before choosing technologies. Decide what should be built internally, what should be borrowed and what simply doesn’t need to exist yet.
  • Shape the technical foundation: As an early infrastructure hire, you’ll have significant influence over architecture, engineering standards and how the research platform develops as the company scales.
What We’re Looking For
  • Strong software engineering fundamentals and excellent coding ability, particularly in Python.
  • Experience building ML infrastructure, ML platforms, MLOps or distributed systems in production.
  • Strong understanding of the lifecycle around modern machine learning systems — data, training, evaluation, deployment, inference and monitoring.
  • Experience working with GPU workloads and an understanding of the performance and reliability challenges surrounding them.
  • Experience with containerisation and orchestration technologies such as Docker and Kubernetes.
  • Strong understanding of cloud infrastructure and Infrastructure-as-Code.
  • Experience designing reliable data pipelines, APIs and distributed services.
  • Ability to reason from first principles about requirements, scale, constraints and trade‑offs, rather than defaulting to technologies you’ve previously used.
  • Comfortable working closely with researchers and translating loosely defined scientific requirements into robust engineering systems.
  • Ability to independently own technically difficult problems in a highly ambiguous environment.
You’ll Likely Thrive Here If
  • You enjoy building infrastructure from first principles rather than inheriting a mature platform with every abstraction already defined.
  • You care about making researchers dramatically more productive.
  • You can move comfortably between ML systems, data engineering, cloud infrastructure and software engineering.
  • You understand that good infrastructure is often about what you choose not to build.
  • You naturally think about GPU utilisation, latency, throughput, reliability, observability and cost.
  • You enjoy profiling systems and finding where the real bottleneck sits.
  • You’re comfortable supporting different models, frameworks and research workflows rather than designing around one narrow use case.
  • You want your infrastructure work to enable scientific discovery and physical‑world outcomes, rather than another consumer or enterprise software product.
  • You enjoy small, highly technical teams where individual engineers have substantial ownership.
Nice To Have
  • Experience with vLLM, TensorRT-LLM, Triton, Ray / KubeRay or similar ML‑serving infrastructure.
  • Experience designing distributed GPU training or inference systems.
  • Experience with PyTorch, JAX or other scientific/deep‑learning frameworks.
  • Experience with large‑scale data ingestion and research‑data platforms.
  • Experience building self‑service ML platforms or tooling for Research Scientists.
  • Experience with model registries, experiment tracking, lineage, evaluation and reproducibility.
  • Strong Kubernetes, Terraform and cloud infrastructure experience.
  • Experience optimising inference through batching, caching, quantisation, scheduling or hardware‑aware optimisation.
  • Experience in AI4Science, scientific computing, HPC, frontier AI or research‑heavy engineering environments.
  • Interest in the intersection of AI, science and automated experimentation.
What’s On Offer
  • Highly competitive compensation + meaningful equity
  • Join a well‑funded, early‑stage AI4Science company at a foundational point in its journey.
  • Significant ownership over the infrastructure underpinning the research organisation.
  • Work alongside exceptional AI researchers, engineers and scientists.
  • Build systems spanning frontier ML, scientific data, simulation and experimentation.
  • Opportunity to influence architecture and engineering culture from an early stage.
  • Work where improvements to infrastructure directly increase the speed at which scientists can experiment, learn and discover.
  • London‑based, collaborative environment with flexibility around hybrid working.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)
Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Good lunch and dinner
Collaborative work culture
No bureaucracy
Principal Machine Learning Infrastructure Engineer
Principal Machine Learning Infrastructure Engineer

PhysicsX • City Of London

On-site
GBP 80,000 - 120,000
Equity options
10% employer pension contribution
Free office lunches
+2
Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)
Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)

Inherentlabs • Greater London

On-site
GBP 70,000 - 90,000
Principal Machine Learning Infrastructure Engineer London, United Kingdom
Principal Machine Learning Infrastructure Engineer London, United Kingdom

PhysicsX Ltd • Greater London

On-site
GBP 80,000 - 100,000
Equity options
10% employer pension contribution
Free office lunches
+6
Senior AI Software Engineer
Senior AI Software Engineer

Ocho • Belfast City District

On-site
GBP 90,000 - 140,000
Share options
Senior Data Scientist
Senior Data Scientist

Ocho • Belfast City District

Hybrid
GBP 90,000 - 140,000
Senior Platform Engineer (Product Initiatives) - Systems Integrator
Senior Platform Engineer (Product Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Remote
GBP 120,000 - 190,000
High-Upside Equity
Flexible remote setup
Work-Life Balance
+1
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Senior MLOps Engineer - AI Infrastructure
Senior MLOps Engineer - AI Infrastructure

Harnham • Greater London

Hybrid
GBP 51,000 - 85,000
Hybrid work model
Engineering Tech Lead - Data & AI
Engineering Tech Lead - Data & AI

The ECA International Group • Greater London

On-site
GBP 80,000 - 120,000
Enhanced Stakeholder Pension Contribution
25 days annual leave
Health, Life Insurance + EAP Wellbeing Support
+9