Founding AI Inference Engineer for Energy AI

Fuse Energy, LLC

Greater London

Hybrid

GBP 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Biannual bonus
Fully expensed tech
Private health insurance
Meal allowance

Job summary

Fuse Energy is building a fully integrated energy company that merges generation, hardware, grid, trading, AI, and distributed energy. We are seeking the founding engineer to own the inference serving layer, guiding performance, reliability, and deployment across a growing cluster.

You will define the serving strategy, design high-throughput workloads, and collaborate with CUDA/GPU teams to optimize throughput and cost per token while meeting ambitious uptime targets.

Qualifications

  • 4+ years building or operating large-scale inference serving systems.
  • Hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding).
  • Strong systems thinking across a large cluster.
  • Ability to work with GPU/CUDA engineers to integrate performance work into a serving system.
  • Track record of making high-stakes architecture calls and owning the outcome.
  • Founding role; comfortable with ambiguity and shaping a new function.

Responsibilities

  • Define Fuse's inference serving strategy and architecture from first principles.
  • Design and build the serving stack: request routing, batching, scheduling and autoscaling for high-throughput, latency-sensitive inference workloads.
  • Own model-level optimisation strategy for serving, applying quantisation, distillation and speculative decoding.
  • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server).
  • Translate throughput, latency and uptime commitments into concrete technical specifications and capacity plans.
  • Act as direct technical owner of inference performance and reliability.
  • Collaborate with CUDA/GPU engineers to integrate custom kernels and hardware performance into the serving layer.
  • Set the standards, tooling and benchmarks this function will run on as it grows.

Skills

Inference serving
Architecture
Batching
KV-cache management
Quantisation
Speculative decoding
System architecture
Autoscaling
GPU collaboration
Serving frameworks

Tools

Triton Inference Server
vLLM
TensorRT-LLM
SGLang
Kubernetes

Job description

Fuse Energy is building a fully integrated energy company that merges generation, hardware, grid, trading, AI, and distributed energy. We are seeking the founding engineer to own the inference serving layer, guiding performance, reliability, and deployment across a growing cluster.

You will define the serving strategy, design high-throughput workloads, and collaborate with CUDA/GPU teams to optimize throughput and cost per token while meeting ambitious uptime targets.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding AI Inference Engineer – Scale & Serving Expert
Founding AI Inference Engineer – Scale & Serving Expert

Fuse Energy • Greater London

On-site
GBP 150,000 - 210,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
AI Inference Engineer
AI Inference Engineer

Fuse Energy, LLC • Greater London

Hybrid
GBP 120,000 - 180,000
Equity
Biannual bonus
Fully expensed tech
+2
AI Inference Engineer
AI Inference Engineer

Fuse Energy • Greater London

On-site
GBP 150,000 - 210,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Lead Backend Engineer – AI-Driven Energy Platform
Lead Backend Engineer – AI-Driven Energy Platform

Fuse Energy • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Equity
Biannual bonus
+3
Founding GPU Engineer - Equity & Biannual Bonus
Founding GPU Engineer - Equity & Biannual Bonus

Fuse Energy • Greater London

On-site
GBP 90,000 - 130,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Head of Engineering — Hands-On, AI & Real-Time Platform
Head of Engineering — Hands-On, AI & Real-Time Platform

Fuse Energy • Greater London

On-site
GBP 180,000 - 250,000
Equity
Biannual bonus
Fully expensed tech
+2
AI Product Owner: Real-World AI for Energy
AI Product Owner: Real-World AI for Energy

Multicoin • Greater London

On-site
GBP 90,000 - 120,000
Biannual bonus scheme
Fully expensed tech to match your need
Private health insurance
+2
AI Operations Manager
AI Operations Manager

Fuse Energy, LLC • Greater London

On-site
GBP 50,000 - 75,000
Competitive salary
Equity sign-on bonus
Biannual bonus scheme
+3
AI Operations Manager
AI Operations Manager

Multicoin • Greater London

On-site
GBP 50,000 - 70,000
Competitive salary and equity sign-on bonus
Biannual bonus scheme
Fully expensed tech
+2
AI Product Owner: Build Real-World AI for Energy
AI Product Owner: Build Real-World AI for Energy

Fuse Energy, LLC • Greater London

Hybrid
GBP 90,000 - 130,000
Equity
Biannual bonus
Tech stipend
+2