Founding GPU Engineer

Fuse Energy

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity sign-on bonus
Biannual bonus
Fully expensed tech
Breakfast and dinner allowance

Job summary

Fuse Energy is seeking a seasoned CUDA performance engineer to design and optimise kernels for high-throughput workloads in a high-performance compute environment.

You will profile GPU bottlenecks, build tooling to align power draw with energy pricing, and optimise multi-GPU scaling across NCCL/MPI. Collaboration with ML and systems teams is essential to improve training and inference pipelines.

Qualifications

  • 4+ years of experience writing production CUDA code or equivalent.
  • Deep understanding of GPU architecture (SMs, warps, memory hierarchy).
  • Proficiency in C++ and CUDA; Python for tooling/automation.
  • Experience with performance profiling tools (Nsight Systems/Compute).
  • Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand).
  • Strong memory optimization, kernel fusion, and parallel algorithm design.
  • Comfortable from kernel to system-level infra.

Responsibilities

  • Design, implement, and optimise CUDA kernels for high-throughput workloads.
  • Profile GPU performance across compute, memory, and interconnect bottlenecks.
  • Build tooling to correlate GPU power draw with energy pricing and grid signals.
  • Optimise multi-GPU/multi-node scaling using NCCL, MPI, or similar libraries.
  • Collaborate with data center teams on power capping and DVFS strategies.
  • Integrate custom kernels into ML training/inference pipelines.
  • Benchmark against CPU/GPU baselines and drive improvements.
  • Contribute to internal libraries, docs, and best practices for GPU performance.

Skills

CUDA programming
C++
Python tooling
GPU architecture knowledge
Performance profiling
Multi-GPU scaling
Memory optimization
Energy-aware compute

Tools

Nsight Systems
NCCL
MPI
RDMA/InfiniBand

Job description

The Opportunity

Demand for high-performance compute capacity across the markets we operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure.

Responsibilities
  • Design, implement, and optimise CUDA kernels for high-throughput, latency-sensitive workloads.
  • Profile and tune GPU performance across compute, memory bandwidth, and interconnect (NVLink/PCIe) bottlenecks.
  • Build tooling to correlate GPU cluster power draw and utilisation with real-time energy pricing and grid signals.
  • Optimise multi-GPU and multi-node scaling using NCCL, MPI, or similar communication libraries.
  • Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity.
  • Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines.
  • Benchmark against CPU/GPU baselines and drive continuous performance improvements.
  • Contribute to internal libraries, documentation, and best practices for GPU performance engineering.
  • 4+ years of experience writing production CUDA code, or equivalent strong project/industry experience.
  • Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy).
  • Proficiency in C++ and CUDA; experience with Python for tooling/orchestration.
  • Experience with performance profiling tools (Nsight Systems/Compute).
  • Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand).
  • Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design.
  • Comfortable working across the stack from low-level kernels to system-level infrastructure.
Nice to Have
  • Experience with Triton, cuDNN, cuBLAS, or custom ML inference/training frameworks.
  • Exposure to data center power/thermal management or demand-response systems.
  • Background in HPC, quantitative finance, or large-scale distributed systems.
  • Familiarity with Kubernetes/Slurm for GPU cluster orchestration.
  • Interest or experience in energy markets, grid systems, or sustainability-focused compute.
  • Competitive salary and an equity sign-on bonus.
  • Biannual bonus scheme.
  • Fully expensed tech to match your needs.
  • Breakfast and dinner allowance for office based employees.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding GPU Engineer
Founding GPU Engineer

Fuse Energy, LLC • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Equity sign-on bonus
Biannual bonus
+2
CUDA Engineer
CUDA Engineer

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity sign-on bonus
Fully expensed tech equipment
Breakfast and dinner allowance
+1
CUDA Engineer
CUDA Engineer

Fuse Energy • Greater London

On-site
GBP 120,000 - 160,000
Equity sign‑on bonus
Biannual bonus
Fully expensed tech
+1
Founding GPU Engineer - Equity & Biannual Bonus
Founding GPU Engineer - Equity & Biannual Bonus

Fuse Energy • Greater London

On-site
GBP 90,000 - 130,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Founding GPU Engineer — CUDA Performance for HPC
Founding GPU Engineer — CUDA Performance for HPC

Fuse Energy, LLC • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Equity sign-on bonus
Biannual bonus
+2
AI Inference Engineer
AI Inference Engineer

Fuse Energy • Greater London

On-site
GBP 150,000 - 210,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
GPU Infrastructure Lead - Systems Integrator
GPU Infrastructure Lead - Systems Integrator

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 140,000 - 170,000
Full Benefits
Senior CUDA Engineer for High-Performance Inference
Senior CUDA Engineer for High-Performance Inference

Fuse Energy • Greater London

On-site
GBP 120,000 - 160,000
Equity sign‑on bonus
Biannual bonus
Fully expensed tech
+1
Technical Solutions Architect – Investors
Technical Solutions Architect – Investors

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000
Senior CUDA Engineer - High-Performance GPU Inference
Senior CUDA Engineer - High-Performance GPU Inference

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity sign-on bonus
Fully expensed tech equipment
Breakfast and dinner allowance
+1