Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited

United Kingdom

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
Full visa sponsorship

Job summary

A leading AI research firm in the UK is seeking a skilled professional to architect AI infrastructure that will redefine engineering through high-speed AI inference. The role includes designing multi-node GPU training environments and collaborating with experts in a vibrant culture. The ideal candidate will be proficient in Kubernetes and experienced in machine learning frameworks. The firm offers competitive equity, a remote-first setup with some in-person presence in London, and full visa sponsorship for top talent.

Qualifications

  • Experience designing multi-node/multi-GPU clusters.
  • Knowledge of optimization techniques like FSDP and custom kernels.
  • Ability to build high-performance pipelines for engineering data.

Responsibilities

  • Design and scale training environments using multiple frameworks.
  • Manage compute strategies across public and sovereign clouds.
  • Implement optimization techniques to maximize performance.
  • Integrate physics solvers into machine learning pipelines.
  • Setup CI/CD and experiment tracking for validation.

Skills

Expert-level Kubernetes (AKS/EKS)
Strong proficiency in PyTorch
Strong proficiency in JAX
Solid experience with Python
Solid experience with Go
Experience with distributed data tools (Dask/Spark)
Background in AI research labs

Tools

NVIDIA NeMo
OpenFOAM
Simcenter

Job description

Looking to architect next-generation AI infrastructure that transforms engineering simulations?

Join a world-class AI research laboratory, currently in Series B moving to C, that is redefining engineering with Large Physical Models that replace traditional numerical simulations with high-speed AI inference. The role involves building and managing the infrastructure required to train complex simulation models on hundreds of GPUs, leading the orchestration of massive-scale training environments, and contributing to pioneering Physical AI architectures that go far beyond standard generative models. Engineers will work in a high-growth, NVIDIA-backed environment with exposure to cutting-edge AI hardware and sovereign cloud technologies.

Ready to design and scale infrastructure that powers the future of AI-driven engineering? Apply now.

Responsibilities
  • Design and scale training environments using PyTorch Distributed, JAX, or NVIDIA NeMo across multi-node/multi-GPU clusters.
  • Manage a mixed-compute strategy spanning public clouds (AWS/Azure) and sovereign industrial clouds for sensitive data.
  • Implement optimization techniques like FSDP and custom kernels to maximize FLOPS for irregular mesh and 3D geometric data.
  • Build high-performance pipelines for ingesting CAE/CFD/FEA engineering data, ensuring zero I/O bottlenecks.
  • Integrate traditional physics solvers (OpenFOAM/Simcenter) into ML pipelines for active learning and model refinement.
  • Setup \"physics-aware\" CI/CD and experiment tracking (Kubeflow/MLFlow) that validates physical consistency laws.
Skills/Must have
  • Orchestration: Expert-level Kubernetes (AKS/EKS) is essential.
  • ML Frameworks: Strong proficiency in PyTorch, JAX, or NVIDIA NeMo.
  • HPC/Data: Solid experience with Python, Go, and distributed data tools (Dask/Spark).
  • Background: Experience in AI research labs (e.g., DeepMind, OpenAI) or Neocloud environments.
Benefits
  • Competitive Equity: Significant stock option packages in a fast-scaling Series B/C firm.
  • Flexible \"London-Plus\" Setup: Remote-first within Europe/UK, with roughly 1 week per month in London (all travel/accommodation fully paid).
  • High-Impact Culture: Work alongside world-renowned physicists, mathematicians, and Formula 1 simulation veterans.
  • Sponsorship: Full visa sponsorship available for top-tier global talent.
Salary
  • Super Competitive (Base + Bonus + High-Upside Equity)
  • Tailored to attract the best in the industry.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Platform Engineer (Product Initiatives) - Systems Integrator
Senior Platform Engineer (Product Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

On-site
GBP 120,000 - 190,000
High-Upside Equity
Flexible remote setup
Work-Life Balance
+1
Principal Machine Learning Infrastructure Engineer
Principal Machine Learning Infrastructure Engineer

PhysicsX • City Of London

On-site
GBP 80,000 - 120,000
Equity options
10% employer pension contribution
Free office lunches
+2
Performance Engineer (Junior) | AI Infrastructure | Cambridge (Hybrid)
Performance Engineer (Junior) | AI Infrastructure | Cambridge (Hybrid)

Pure Resourcing Solutions Limited • Linton

On-site
GBP 42,000 - 70,000
Pension
Hybrid from Cambridge office
Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid)
Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid)

Pure Resourcing Solutions Limited • Linton

On-site
GBP 90,000 - 120,000
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

remotestar-team • Cambourne

Hybrid
GBP 90,000 - 150,000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+7
Senior Machine Learning Software Engineer, Research
Senior Machine Learning Software Engineer, Research

PhysicsX • City Of London

On-site
GBP 60,000 - 80,000
10% employer pension contribution
Private medical insurance
25 days Annual Leave
+2
Member of Technical Staff - Research Software Engineer
Member of Technical Staff - Research Software Engineer

Reflection • Greater London

On-site
GBP 70,000 - 100,000
Top-tier compensation
Comprehensive medical, dental, vision insurance
Fully paid parental leave
+2
Research Engineer - Data Infrastructure/ML
Research Engineer - Data Infrastructure/ML

Thirddimension • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary & stock options
Pension / retirement plan
Health & wellness
+5
Performance Engineer (Junior) | AI Infrastructure | Cambridge
Performance Engineer (Junior) | AI Infrastructure | Cambridge

Pure Resourcing Solutions • Dry Drayton

On-site
GBP 63,000 - 77,000
Hybrid work model
Machine Learning Engineer
Machine Learning Engineer

Amberes Recruitment • Greater London

On-site
GBP 85,000 - 110,000
Bonus
Benefits
Equity