CUDA Engineer

Fuse Energy Supply

Greater London

On-site

GBP 63,000 - 103,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity sign-on bonus
Biannual bonus
Tech stipend
Breakfast and dinner allowance

Job summary

Fuse Energy Supply in the United Kingdom is seeking an experienced CUDA specialist to write and optimize production kernels for transformer inference workloads, with a focus on latency and throughput.

You will profile and tune kernels, apply fusion, manage memory hierarchies, and implement mixed-precision and quantisation-aware code. Experience with multi-GPU, high-performance libraries, and reading PTX/SASS is a plus.

Qualifications

  • 4+ years writing production CUDA code with shipping performance-critical kernels.
  • Deep understanding of GPU microarchitecture, warps, occupancy, and memory hierarchy.
  • Strong CUDA C++ skills, including streams and asynchronous execution.
  • Hands-on profiling to diagnose compute-bound vs. memory-bound bottlenecks.
  • Experience with kernel fusion, memory coalescing, and avoiding warp divergence.
  • Experience writing quantised and mixed-precision kernels.
  • Solid grasp of parallel algorithm design and numerical precision tradeoffs.
  • Nice to have: transformer/attention-style kernels or autoregressive decoding.
  • Nice to have: building high-performance GPU libraries from scratch.
  • Nice to have: background in HPC or latency-critical perf engineering.
  • Nice to have: exposure to multi-GPU or multi-node kernel-level optimisation.
  • Nice to have: comfortable reading PTX/SASS to validate kernel efficiency.

Responsibilities

  • Write and optimise custom CUDA kernels for core transformer inference operations.
  • Profile kernels to identify and eliminate bottlenecks in occupancy, memory throughput, and warp divergence.
  • Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines.
  • Optimise memory access patterns and manage the memory hierarchy for maximum bandwidth utilisation.
  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprint.
  • Build and tune caching mechanisms for efficient autoregressive decoding.
  • Tune kernel launch configurations for target GPU architectures.
  • Benchmark kernels against existing baselines and drive measurable throughput and latency improvements.
  • Write tests for CUDA code to catch performance and correctness regressions.
  • Maintain internal CUDA libraries and contribute to our team coding standards and documentation.

Skills

CUDA
CUDA C++
GPU microarchitecture
Profiling
Kernel fusion
Memory coalescing
Warp divergence
Quantised kernels
Mixed-precision
Parallel algorithms
Transformer kernels
High-performance libraries
HPC experience
Multi-GPU knowledge
PTX/SASS reading

Tools

SASS
NodeJS
CUDA

Job description

Salary: £63,000 - 103,000 per year

Requirements:
  • 4+ years writing production CUDA code, with a track record of shipping performance-critical kernels
  • Deep understanding of GPU microarchitecture, warps, occupancy, register pressure, and memory hierarchy
  • Strong CUDA C++ skills, including streams and asynchronous execution
  • Hands-on experience profiling to diagnose compute-bound vs. memory-bound bottlenecks
  • Experience with kernel fusion, memory coalescing, and avoiding warp divergence
  • Experience writing quantised and mixed-precision kernels
  • Solid grasp of parallel algorithm design and numerical precision tradeoffs
  • Nice to have: experience with transformer/attention-style kernels or autoregressive decoding
  • Nice to have: experience building high-performance GPU libraries from scratch
  • Nice to have: background in HPC or other latency-critical performance engineering
  • Nice to have: exposure to multi-GPU or multi-node kernel-level optimisation
  • Nice to have: comfortable reading PTX/SASS to validate kernel efficiency
Responsibilities:
  • Write and optimise custom CUDA kernels for core transformer inference operations
  • Profile kernels to identify and eliminate bottlenecks in occupancy, memory throughput, and warp divergence
  • Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines
  • Optimise memory access patterns and manage the memory hierarchy for maximum bandwidth utilisation
  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprint
  • Build and tune caching mechanisms for efficient autoregressive decoding
  • Tune kernel launch configurations for target GPU architectures
  • Benchmark kernels against existing baselines and drive measurable throughput and latency improvements
  • Write tests for CUDA code to catch performance and correctness regressions
  • Maintain internal CUDA libraries and contribute to our team coding standards and documentation
Technologies:
  • AI
  • CUDA
  • SASS
  • NodeJS
More:

We are Fuse Energy, a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system. Backed by $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels, we are expanding into high-performance compute infrastructure at the intersection of energy and AI. We focus on optimising how power-dense GPU workloads are scheduled, cooled, and balanced against grid conditions in real time. This is a full-time role with competitive salary, an equity sign-on bonus, a biannual bonus scheme, fully expensed tech to match your needs, and a breakfast and dinner allowance for office-based employees.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Engineer
CUDA Engineer

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity sign-on bonus
Fully expensed tech equipment
Breakfast and dinner allowance
+1
CUDA Engineer
CUDA Engineer

Fuse Energy • Greater London

On-site
GBP 120,000 - 160,000
Equity sign‑on bonus
Biannual bonus
Fully expensed tech
+1
Founding GPU Engineer
Founding GPU Engineer

Fuse Energy, LLC • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Equity sign-on bonus
Biannual bonus
+2
Founding GPU Engineer
Founding GPU Engineer

Fuse Energy • Greater London

On-site
GBP 90,000 - 130,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Senior CUDA Engineer for High-Performance Inference
Senior CUDA Engineer for High-Performance Inference

Fuse Energy • Greater London

On-site
GBP 120,000 - 160,000
Equity sign‑on bonus
Biannual bonus
Fully expensed tech
+1
AI Inference Engineer
AI Inference Engineer

Fuse Energy Supply • Greater London

On-site
GBP 63,000 - 103,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
CUDA Kernel Engineer - High-Performance Transformer Inference
CUDA Kernel Engineer - High-Performance Transformer Inference

Fuse Energy Supply • Greater London

On-site
GBP 63,000 - 103,000
Equity sign-on bonus
Biannual bonus
Tech stipend
+1
Senior CUDA Engineer - High-Performance GPU Inference
Senior CUDA Engineer - High-Performance GPU Inference

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity sign-on bonus
Fully expensed tech equipment
Breakfast and dinner allowance
+1
Founding GPU Engineer - Equity & Biannual Bonus
Founding GPU Engineer - Equity & Biannual Bonus

Fuse Energy • Greater London

On-site
GBP 90,000 - 130,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
AI Inference Engineer
AI Inference Engineer

Fuse Energy • Greater London

On-site
GBP 150,000 - 210,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1