Staff ML Performance Engineer (Compiler)

EngineersOfAI

Greater London

On-site

GBP 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Wayve in London is seeking a Staff ML Performance Engineer to optimise inference on edge accelerators and GPUs for our first driving product. You will shape the technical direction for production systems running on in-vehicle compute, spanning ML systems, compilers, runtimes and embedded deployment.

This role combines hands-on work with strategy across multiple early-stage projects. The ideal candidate will have deep experience optimizing production ML workloads under tight constraints and be

Qualifications

  • Proven experience improving performance in production systems with tight constraints.
  • Strong proficiency with ML stacks and inference toolchains; ability to learn adjacent frameworks quickly.
  • Comfort operating across high-level models to low-level kernel/runtime execution.
  • Solid software engineering fundamentals and debugging expertise.
  • Excellent communication and collaboration across multiple stakeholders.

Responsibilities

  • Identify, implement and validate optimisations in compilers, runtimes, and kernels.
  • Profile bottlenecks across the full inference stack and deliver measurable improvements.
  • Build benchmarking and regression tests to ensure performance across models and devices.
  • Develop and optimise for multiple target platforms (e.g., NVIDIA Orin/Thor, Qualcomm).
  • Collaborate with model developers to influence architecture and deployment decisions.
  • Contribute to roadmaps and tooling to raise performance engineering standards.

Skills

Performance optimization
ML inference
Edge deployment
CUDA/TensorRT
C++/Python
Embedded systems

Tools

TensorRT
CUDA
Qualcomm QNN
Triton
OpenCL
MLIR

Job description

About us

Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.
In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.

At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.

Make Wayve the experience that defines your career!

The role

As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and GPUs. The focus of this team is to run large transformer-based models efficiently on low-cost, low-power edge devices to enable Wayve’s first driving product.

You’ll help set the technical direction for turning these models into production systems that run reliably on in-vehicle compute. This is a hands‑on role working across ML systems, compilers, runtimes, kernels, and embedded deployment, contributing to several early-stage, high-impact projects at Wayve.

Key responsibilities
  • Identify, implement and validate optimisations in ML compilers, runtimes, and kernels (e.g. operator fusion, scheduling, quantisation‑aware performance, custom kernels)
  • Profile and pinpoint bottlenecks across the full inference stack (model graph, compiler/runtime, kernel execution, memory movement) and deliver measurable improvements.
  • Build robust benchmarking and regression testing to ensure performance improvements hold across models, devices, and software releases.
  • Develop and optimise for multiple target platforms (e.g. NVIDIA Orin/Thor, Qualcomm), working with cross‑functional teams to deliver performant and maintainable solutions.
  • Collaborate with model developers to influence architecture and training/deployment decisions that affect on‑device performance.
  • Contribute to technical roadmaps and tooling and help raise the standard of performance engineering across the team
About you
Essential
  • Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost).
  • Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, MLIR, ONNX) and confidence learning adjacent frameworks quickly.
  • Comfort operating at multiple levels of abstraction — from high‑level model behaviour down to low‑level kernel/runtime execution.
  • Strong software engineering fundamentals (debugging, profiling, testing, and maintainable code).
  • Clear communicator and collaborative teammate; able to align multiple stakeholders on performance trade‑offs and priorities.
Desirable
  • Experience with compute graph scheduling and execution on multiple targets
  • Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system‑level constraints.
  • Experience with NVIDIA and/or Qualcomm SoCs and performance tooling.
  • Python and C++ proficiency.
  • Experience mentoring others and/or driving technical direction in a small, fast‑moving team.

#LI-HH1

Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know.
We understand that everyone has a unique set of skil

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Performance Engineer (Compiler)
Staff ML Performance Engineer (Compiler)

Wayve • Greater London

On-site
GBP 120,000 - 180,000
Staff ML Performance Engineer (Compiler)
Staff ML Performance Engineer (Compiler)

Icehouseventures • Greater London

On-site
GBP 90,000 - 130,000
Staff ML Performance Engineer (Compiler)
Staff ML Performance Engineer (Compiler)

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 170,000
Staff ML Performance Engineer (Inference Optimisation)
Staff ML Performance Engineer (Inference Optimisation)

Wayve • Greater London

On-site
GBP 70,000 - 90,000
Senior Machine Learning Engineer, AI Performance
Senior Machine Learning Engineer, AI Performance

EngineersOfAI • Greater London

On-site
GBP 90,000 - 130,000
Senior Machine Learning Engineer, AI Performance
Senior Machine Learning Engineer, AI Performance

Wayve • Greater London

On-site
GBP 90,000 - 130,000
Senior Machine Learning Engineer, AI Performance
Senior Machine Learning Engineer, AI Performance

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 140,000
Staff ML Performance Engineer - Edge Inference Optimizer
Staff ML Performance Engineer - Edge Inference Optimizer

Wayve • Greater London

On-site
GBP 120,000 - 180,000
Staff ML Performance Engineer: Edge Inference & Systems
Staff ML Performance Engineer: Edge Inference & Systems

AI Startups UK • Greater London

Hybrid
GBP 120,000 - 170,000
Staff ML Performance Engineer: Edge Inference Optimizer
Staff ML Performance Engineer: Edge Inference Optimizer

EngineersOfAI • Greater London

Hybrid
GBP 120,000 - 180,000