AI Inference Engineer

Fuse Energy Supply

Greater London

On-site

GBP 63,000 - 103,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity sign-on bonus
Biannual bonus
Fully expensed tech
Breakfast and dinner allowance

Job summary

Fuse Energy seeks a founding engineer to define and build our inference serving stack, driving architecture decisions for high-throughput, low-latency workloads. You will own model-level optimisation strategies, including quantisation and speculative decoding, and collaborate with CUDA/GPU teams to integrate performance work.

The role reports directly to the CTO in a pioneering energy-AI platform, with equity and competitive compensation in a fast-growing startup environment.

Qualifications

  • 4+ years of experience building or operating large-scale inference serving systems.

Responsibilities

  • Define our inference serving strategy and architecture from first principles.

Skills

Inference serving
GPU/CUDA collaboration
System architecture
Performance optimisation
Kubernetes
Multi-tenant infra
Hyperscaler experience
Quantisation/Distillation
SLA-driven infra
Leadership

Tools

vLLM
TensorRT-LLM
Triton Inference Server
SGLang

Job description

Salary: £63,000 - 103,000 per year

Requirements:
  • 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them, including batching, KV-cache management, quantisation, and speculative decoding.
  • Strong systems thinking, with the ability to reason about the full path from incoming request to served response across a large cluster.
  • Comfort working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
  • A track record of making high-stakes architecture calls and owning the outcome.
  • Comfort operating without a playbook in a founding role shaping a new function around early-stage architecture.
  • Experience with Triton or custom ML inference/training frameworks is nice to have.
  • Experience with autoscaling or capacity planning for large-scale inference workloads is nice to have.
  • Exposure to multi-tenant serving or SLA-driven infrastructure is nice to have.
  • Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system is nice to have.
  • Familiarity with Kubernetes/Slurm for cluster orchestration is nice to have.
  • Interest or experience in energy markets, grid systems, or sustainability-focused compute is nice to have.
Responsibilities:
  • Define our inference serving strategy and architecture from first principles.
  • Design and build the serving stack, including request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
  • Own our model-level optimisation strategy for serving, deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with our CUDA/GPU engineers.
  • Make the core software architecture decisions on serving frameworks and orchestration, such as vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents.
  • Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
  • Act as a direct technical owner of inference performance and reliability.
  • Work closely with our CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
  • Set the standards, tooling, and benchmarks this function will run on as it grows.
Technologies:
  • AI
  • CTO
  • CUDA
  • Hardware
  • Kubernetes
  • LLM
  • vLLM
More:

We are Fuse Energy, a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system, and we have raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group, and strategic angels such as Nico Rosberg, the Co-Founder of Solana, and GPs behind Meta, Revolut, Spotify, Uber, and more. As data centres become one of the largest and fastest-growing sources of electricity demand, we are expanding into high-performance compute infrastructure at the intersection of energy and AI. We are building the GPU/CUDA performance layer and the inference serving layer from scratch, and this founding engineer role reports directly to our CTO. We offer a competitive salary, an equity sign-on bonus, a biannual bonus scheme, fully expensed tech to match your needs, and a breakfast and dinner allowance for office-based employees.

  • A competitive salary
  • An equity sign-on bonus
  • A biannual bonus scheme
  • Fully expensed tech to match your needs
  • A breakfast and dinner allowance for office-based employees

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

Fuse Energy • Greater London

On-site
GBP 150,000 - 210,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
AI Inference Engineer — Equity Sign-On + High-Throughput Serving
AI Inference Engineer — Equity Sign-On + High-Throughput Serving

Fuse Energy Supply • Greater London

On-site
GBP 63,000 - 103,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
CUDA Engineer
CUDA Engineer

Fuse Energy Supply • Greater London

On-site
GBP 63,000 - 103,000
Equity sign-on bonus
Biannual bonus
Tech stipend
+1
Founding AI Inference Engineer – Scale & Serving Expert
Founding AI Inference Engineer – Scale & Serving Expert

Fuse Energy • Greater London

On-site
GBP 150,000 - 210,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

On-site
GBP 70,000 - 95,000
Equity options
Competitive compensation
Staff Software Engineer, Inference
Staff Software Engineer, Inference

CoreWeave • Greater London

On-site
GBP 120,000 - 180,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+6
CUDA Engineer
CUDA Engineer

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity sign-on bonus
Fully expensed tech equipment
Breakfast and dinner allowance
+1
CUDA Engineer
CUDA Engineer

Fuse Energy • Greater London

On-site
GBP 120,000 - 160,000
Equity sign‑on bonus
Biannual bonus
Fully expensed tech
+1
AI Inference Engineer | GPU-Scale Rust/Python | Equity
AI Inference Engineer | GPU-Scale Rust/Python | Equity

Perplexity • Greater London

On-site
GBP 70,000 - 95,000
Lead AI Architect
Lead AI Architect

Intellectual Capital Resources • Cambridge

On-site
GBP 59,000 - 99,000
Pension
Health benefits