AI Inference Engineer

Fuse Energy

Greater London

On-site

GBP 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity sign-on bonus
Biannual bonus
Fully expensed tech
Breakfast and dinner allowance

Job summary

Fuse Energy is seeking a founding AI Inference Engineer to define and build how we serve AI workloads at scale, reporting to the CTO. You’ll own the layer above CUDA/GPU work, shaping inference delivery, throughput, and reliability as we scale data-centre compute for energy applications.

With 4+ years in large-scale inference systems, you’ll collaborate with GPU engineers on model-level optimisations and architecture decisions, driving high-impact solutions across the platform.

Qualifications

  • 4+ years of experience building or operating large-scale inference serving systems.
  • Strong understanding of inference serving frameworks and optimisation techniques (batching, KV-cache management, quantisation, speculative decoding).
  • Ability to reason about full request-to-response path across a large cluster.
  • Track record of making high-stakes architecture calls and owning outcomes.

Responsibilities

  • Define inference serving strategy and architecture from first principles.
  • Design and build the serving stack: routing, batching, scheduling, autoscaling for high-throughput, low-latency workloads.
  • Own model-level optimisation strategy to improve throughput and cost per token, in collaboration with CUDA/GPU engineers.
  • Make core architectural calls on serving frameworks and orchestration (e.g., vLLM, TensorRT-LLM, SGLang, Triton).
  • Translate throughput, latency, and uptime commitments into concrete specifications and capacity plans.
  • Act as a direct technical owner of inference performance and reliability.
  • Collaborate with CUDA and GPU teams to integrate performance work into the serving layer.
  • Set standards, tooling, and benchmarks for this function as it grows.

Skills

Inference serving
System architecture
GPU/CUDA integration
Performance optimisation
Founding role experience

Tools

Triton Inference Server
vLLM
TensorRT-LLM
SGLang

Job description

Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy -- fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch -- and we're looking for the founding engineer to own the latter.

We're looking for a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered against committed performance targets.

The Opportunity

Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market. That's this role.

Responsibilities
  • Define Fuse's inference serving strategy and architecture from first principles.
  • Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
  • Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers.
  • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).
  • Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
  • Act as a direct technical owner of inference performance and reliability.
  • Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
  • Set the standards, tooling, and benchmarks this function will run on as it grows.
Qualifications
  • 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding).
  • Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster.
  • Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
  • A track record of making high-stakes architecture calls and owning the outcome.
  • Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one.
Nice to Have
  • Experience with Triton or custom ML inference/training frameworks.
  • Experience with autoscaling or capacity planning for large-scale inference workloads.
  • Exposure to multi-tenant serving or SLA-driven infrastructure.
  • Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
  • Familiarity with Kubernetes/Slurm for cluster orchestration.
  • Interest or experience in energy markets, grid systems, or sustainability-focused compute.
Benefits
  • Competitive salary and an equity sign-on bonus.
  • Biannual bonus scheme.
  • Fully expensed tech to match your needs.
  • Breakfast and dinner allowance for office based employees.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Engineer
CUDA Engineer

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity sign-on bonus
Fully expensed tech equipment
Breakfast and dinner allowance
+1
CUDA Engineer
CUDA Engineer

Fuse Energy • Greater London

On-site
GBP 120,000 - 160,000
Equity sign‑on bonus
Biannual bonus
Fully expensed tech
+1
Founding AI Inference Engineer – Scale & Serving Expert
Founding AI Inference Engineer – Scale & Serving Expert

Fuse Energy • Greater London

On-site
GBP 150,000 - 210,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Founding GPU Engineer
Founding GPU Engineer

Fuse Energy • Greater London

On-site
GBP 90,000 - 130,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Founding GPU Engineer
Founding GPU Engineer

Fuse Energy, LLC • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Equity sign-on bonus
Biannual bonus
+2
AI Operations Manager
AI Operations Manager

Fuse Energy, LLC • Greater London

On-site
GBP 50,000 - 75,000
Competitive salary
Equity sign-on bonus
Biannual bonus scheme
+3
AI Operations Manager
AI Operations Manager

Multicoin • Greater London

On-site
GBP 50,000 - 70,000
Competitive salary and equity sign-on bonus
Biannual bonus scheme
Fully expensed tech
+2
Applied AI Engineer
Applied AI Engineer

Fuse Energy • Greater London

On-site
GBP 70,000 - 90,000
Competitive salary
Stock options sign-on bonus
Biannual bonus scheme
+3
Senior CUDA Engineer for High-Performance Inference
Senior CUDA Engineer for High-Performance Inference

Fuse Energy • Greater London

On-site
GBP 120,000 - 160,000
Equity sign‑on bonus
Biannual bonus
Fully expensed tech
+1
M&A Strategy Associate
M&A Strategy Associate

Fuse Energy • Greater London

Hybrid
GBP 60,000 - 80,000
Competitive salary and equity sign-on bonus
Bi-annual bonus scheme
Fully expensed tech
+3