Senior ML Systems Engineer - Simulations

Oriole

Greater London

On-site

GBP 90,000 - 150,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Oriole Networks, based in London, seeks a Senior ML Systems Engineer to build and validate simulation infrastructure for large-scale ML systems. You will model compute, memory, interconnect, and communication behavior to guide architecture, performance optimization, and capacity planning.

Responsibilities include running benchmarks on real ML workloads, calibrating models, and collaborating with hardware, software, networking, and ML teams to align simulations with actual workloads and

Qualifications

  • Master's or PhD in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • Strong experience in ML systems, distributed systems, performance engineering, computer architecture, or simulation.
  • Understanding of systems used for machine learning training and inference.
  • Experience analyzing compute, communication, and memory behavior in large-scale ML systems.
  • Hands-on experience with performance benchmarking, profiling, and measurement of ML systems.
  • Experience with distributed training concepts such as data parallelism, tensor/model parallelism, pipeline parallelism, collectives, and synchronization overheads.
  • Proficiency in Python, C++, or Rust.
  • Strong analytical skills and the ability to connect simulation results to real system behavior.

Responsibilities

  • Build simulation models for compute, memory, interconnect, and communication behavior in ML systems.
  • Develop tools to simulate performance for training and inference workloads.
  • Model distributed execution across accelerators, hosts, and network fabrics, including collectives, synchronization, and communication bottlenecks.
  • Use simulation and analytical modelling to evaluate tradeoffs, identify bottlenecks, and guide system design.
  • Run performance experiments and benchmarks on real ML systems to calibrate and validate simulation models.
  • Analyze end-to-end performance, including throughput, latency, scaling efficiency, utilization, and cost/performance tradeoffs.
  • Partner with hardware/software/Networking/ML teams to align simulation with real workloads and constraints.
  • Create reproducible benchmarking methodologies across models, system configurations, and compare against real system measurements to prove validity.
  • Communicate findings through technical reports and design recommendations.

Skills

ML systems
Distributed systems
Performance engineering
System modeling
Analytical skills

Education

Master's/PhD

Tools

Python
C++
Rust

Job description

We are looking for a Senior ML Systems Engineer to build and validate simulation infrastructure for large-scale machine learning systems. This role focuses on modelling the compute and communication behaviour of systems used for ML training and inference, and using simulation to guide architecture, performance optimization, and capacity planning.

What You'll Do
  • Build simulation models for compute, memory, interconnect, and communication behavior in ML systems.
  • Develop tools to simulate performance for training and inference workloads.
  • Model distributed execution across accelerators, hosts, and network fabrics, including collectives, synchronization, and communication bottlenecks.
  • Use simulation and analytical modelling to evaluate tradeoffs, identify bottlenecks, and guide system design.
  • Run performance experiments and benchmarks on real ML systems to calibrate and validate simulation models.
  • Analyze end-to-end performance, including throughput, latency, scaling efficiency, utilization, and cost/performance tradeoffs.
  • Partner with hardware/software/Networking/ML teams to align simulation with real workloads and constraints.
  • Create reproducible benchmarking methodologies across models, system configurations, and compare against real system measurements to prove validity.
  • Communicate findings through technical reports and design recommendations.
Qualifications
Required
  • Master's, or PhD in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • Strong experience in ML systems, distributed systems, performance engineering, computer architecture, or simulation.
  • Understanding of systems used for machine learning training and inference.
  • Experience analyzing compute, communication, and memory behavior in large-scale ML systems.
  • Hands-on experience with performance benchmarking, profiling, and measurement of ML systems.
  • Experience with distributed training concepts such as data parallelism, tensor/model parallelism, pipeline parallelism, collectives, and synchronization overheads.
  • Proficiency in one of the following Python, C++, or Rust.
  • Strong analytical skills and the ability to connect simulation results to real system behavior.
Preferred
  • Experience with system performance modelling, network simulation, or architecture evaluation tools. - this background is ideal
  • Familiarity with accelerator-based systems such as GPUs, TPUs, or custom ML hardware.
  • Experience with PyTorch, JAX, TensorFlow, NCCL, XLA, CUDA, or similar tools.
  • Knowledge of interconnect and networking technologies such as InfiniBand, Ethernet/RDMA, NVLink, PCIe, or equivalent.
  • Experience evaluating both training throughput and inference latency/serving efficiency.
  • Background in workload characterization, trace-driven simulation, or model calibration.
  • Ability to work across hardware and software boundaries in a cross-functional environment.
What Success Looks Like
  • Build simulation models that accurately predict performance trends and inform architectural decisions.
  • Identify compute and communication bottlenecks in ML training and inference systems.
  • Correlate simulation outputs with real-world benchmark data.
  • Improve system efficiency, scalability, and cost effectiveness through data-driven insights.

Accelerating AI in a Low Carbon World - Oriole Networks is a photonic networking company, developing disruptive technologies for AI/ML and HPC networking that will revolutionise data centres.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Simulation & Performance Architect
Senior ML Systems Simulation & Performance Architect

Oriole • Greater London

On-site
GBP 90,000 - 150,000
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Hybrid
GBP 90,000 - 130,000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Quant Blueprint LLC • Greater London

On-site
GBP 50,000 - 70,000
Senior Machine Learning Engineer (Large Systems)
Senior Machine Learning Engineer (Large Systems)

EngineersOfAI • Cambridge

On-site
GBP 110,000 - 140,000
Senior Machine Learning Engineer (Large Systems)
Senior Machine Learning Engineer (Large Systems)

EngineersOfAI • Bristol

On-site
GBP 90,000 - 140,000
Senior ML Systems Engineer — C++ / PyTorch & Performance
Senior ML Systems Engineer — C++ / PyTorch & Performance

CamWebDir • United Kingdom

Hybrid
GBP 75,000 - 110,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Trading Interview • Greater London

Hybrid
GBP 120,000 - 180,000
Senior Machine Learning Engineer (Large Systems)
Senior Machine Learning Engineer (Large Systems)

EngineersOfAI • Greater London

On-site
GBP 110,000 - 140,000
Principal Machine Learning Infrastructure Engineer London, United Kingdom
Principal Machine Learning Infrastructure Engineer London, United Kingdom

PhysicsX Ltd • Greater London

On-site
GBP 80,000 - 100,000
Equity options
10% employer pension contribution
Free office lunches
+6