Founding AI Inference Engineer – Scale & Serving Expert

Fuse Energy

Greater London

On-site

GBP 150,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity sign-on bonus
Biannual bonus
Fully expensed tech
Breakfast and dinner allowance

Job summary

Fuse Energy is seeking a founding AI Inference Engineer to define and build how we serve AI workloads at scale, reporting to the CTO. You’ll own the layer above CUDA/GPU work, shaping inference delivery, throughput, and reliability as we scale data-centre compute for energy applications.

With 4+ years in large-scale inference systems, you’ll collaborate with GPU engineers on model-level optimisations and architecture decisions, driving high-impact solutions across the platform.

Qualifications

  • 4+ years of experience building or operating large-scale inference serving systems.
  • Strong understanding of inference serving frameworks and optimisation techniques (batching, KV-cache management, quantisation, speculative decoding).
  • Ability to reason about full request-to-response path across a large cluster.
  • Track record of making high-stakes architecture calls and owning outcomes.

Responsibilities

  • Define inference serving strategy and architecture from first principles.
  • Design and build the serving stack: routing, batching, scheduling, autoscaling for high-throughput, low-latency workloads.
  • Own model-level optimisation strategy to improve throughput and cost per token, in collaboration with CUDA/GPU engineers.
  • Make core architectural calls on serving frameworks and orchestration (e.g., vLLM, TensorRT-LLM, SGLang, Triton).
  • Translate throughput, latency, and uptime commitments into concrete specifications and capacity plans.
  • Act as a direct technical owner of inference performance and reliability.
  • Collaborate with CUDA and GPU teams to integrate performance work into the serving layer.
  • Set standards, tooling, and benchmarks for this function as it grows.

Skills

Inference serving
System architecture
GPU/CUDA integration
Performance optimisation
Founding role experience

Tools

Triton Inference Server
vLLM
TensorRT-LLM
SGLang

Job description

Fuse Energy is seeking a founding AI Inference Engineer to define and build how we serve AI workloads at scale, reporting to the CTO. You’ll own the layer above CUDA/GPU work, shaping inference delivery, throughput, and reliability as we scale data-centre compute for energy applications.

With 4+ years in large-scale inference systems, you’ll collaborate with GPU engineers on model-level optimisations and architecture decisions, driving high-impact solutions across the platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

Fuse Energy • Greater London

On-site
GBP 150,000 - 210,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Senior CUDA Engineer for High-Performance Inference
Senior CUDA Engineer for High-Performance Inference

Fuse Energy • Greater London

On-site
GBP 120,000 - 160,000
Equity sign‑on bonus
Biannual bonus
Fully expensed tech
+1
Founding GPU Engineer - Equity & Biannual Bonus
Founding GPU Engineer - Equity & Biannual Bonus

Fuse Energy • Greater London

On-site
GBP 90,000 - 130,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Senior CUDA Engineer - High-Performance GPU Inference
Senior CUDA Engineer - High-Performance GPU Inference

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity sign-on bonus
Fully expensed tech equipment
Breakfast and dinner allowance
+1
Platform Engineer – Scale GPU Infra for AI Platform
Platform Engineer – Scale GPU Infra for AI Platform

Ineffable Intelligence LTD • Greater London

Hybrid
GBP 85,000 - 120,000
Founding GPU Engineer — CUDA Performance for HPC
Founding GPU Engineer — CUDA Performance for HPC

Fuse Energy, LLC • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Equity sign-on bonus
Biannual bonus
+2
Platform Engineer – Scale AI Infra & GPU Orchestration
Platform Engineer – Scale AI Infra & GPU Orchestration

Ineffable Intelligence • Greater London

On-site
GBP 70,000 - 110,000
CUDA Engineer
CUDA Engineer

Fuse Energy • Greater London

On-site
GBP 120,000 - 160,000
Equity sign‑on bonus
Biannual bonus
Fully expensed tech
+1
CUDA Engineer
CUDA Engineer

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity sign-on bonus
Fully expensed tech equipment
Breakfast and dinner allowance
+1
Senior Data Platform Engineer — AI-Powered Systems
Senior Data Platform Engineer — AI-Powered Systems

Scale AI • Greater London

On-site
GBP 110,000 - 170,000