CUDA Engineer

Fuse Energy

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Biannual bonus scheme
Fully expensed tech
Private health insurance
Breakfast and dinner allowance

Job summary

Fuse Energy is hiring a CUDA Engineer to write and optimise low-level GPU code for transformer inference workloads. You will design custom CUDA kernels, tune memory and compute bottlenecks, and push throughput across our fleet.

You have 4+ years in production CUDA, deep GPU microarchitecture knowledge, and experience with kernel fusion, mixed precision, and autoregressive decoding. We offer competitive salary, equity, and private health benefits.

Qualifications

  • 4+ years writing production CUDA code with performance kernels.
  • Deep understanding of GPU microarchitecture: warps, occupancy, registers.
  • Strong CUDA C++ skills with streams and asynchronous execution.
  • Hands-on profiling to diagnose compute-bound vs memory-bound bottlenecks.
  • Experience with kernel fusion, memory coalescing, warp-divergence avoidance.
  • Experience writing quantised and mixed-precision kernels.
  • Knowledge of parallel algorithms and numerical precision tradeoffs.
  • Bonus: transformer/attention-style kernels or autoregressive decoding; multi-GPU or multi-node kernel-level optimisation.

Responsibilities

  • Write and optimise custom CUDA kernels for core transformer inference operations
  • Profile kernels to identify bottlenecks in occupancy, memory throughput, and warp divergence
  • Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines
  • Optimise memory access patterns and manage the memory hierarchy for maximum bandwidth utilisation
  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprint
  • Build and tune caching mechanisms for efficient autoregressive decoding
  • Tune kernel launch configurations for target GPU architectures
  • Benchmark kernels against baselines and drive throughput and latency improvements
  • Write tests for CUDA code to catch performance and correctness regressions
  • Maintain internal CUDA libraries and contribute to coding standards and docs
  • 4+ years writing production CUDA code, with track record of shipping performance-critical kernels
  • Deep understanding of GPU microarchitecture: warps, occupancy, register pressure and memory hierarchy
  • Strong CUDA C++ skills, including streams and asynchronous execution
  • Hands-on experience profiling to diagnose compute-bound vs memory-bound bottlenecks
  • Experience with kernel fusion, memory coalescing and avoiding warp divergence
  • Experience writing quantised and mixed-precision kernels
  • Solid grasp of parallel algorithm design and numerical precision tradeoffs
  • Bonus: transformer/attention-style kernels or autoregressive decoding; building high-performance GPU libraries from scratch; HPC or latency-critical performance engineering; multi-GPU or multi-node kernel-level optimisation

Skills

CUDA C++
Kernel profiling
Kernel fusion
Memory optimization
Quantised kernels
Mixed precision
Autoregressive decoding
Multi-GPU kernels
Latency optimization

Job description

Fuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.

We’ve raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.

We’re building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We’re looking for a CUDA Engineer to write and optimise the low-level GPU code that powers our inference workloads: designing custom CUDA kernels, tuning performance across memory bandwidth and compute bottlenecks, and squeezing maximum throughput out of every GPU in our fleet, working at the level of SMs, warps and memory hierarchies.

Responsibilities
  • Write and optimise custom CUDA kernels for core transformer inference operations

  • Profile kernels to identify and eliminate bottlenecks in occupancy, memory throughput, and warp divergence

  • Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines

  • Optimise memory access patterns and manage the memory hierarchy for maximum bandwidth utilisation

  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprint

  • Build and tune caching mechanisms for efficient autoregressive decoding

  • Tune kernel launch configurations for target GPU architectures

  • Benchmark kernels against existing baselines and drive measurable throughput and latency improvements

  • Write tests for CUDA code to catch performance and correctness regressions

  • Maintain internal CUDA libraries and contribute to team coding standards and documentation

  • 4+ years writing production CUDA code, with a track record of shipping performance-critical kernels

  • Deep understanding of GPU microarchitecture: warps, occupancy, register pressure and memory hierarchy

  • Strong CUDA C++ skills, including streams and asynchronous execution

  • Hands-on experience profiling to diagnose compute-bound vs memory-bound bottlenecks

  • Experience with kernel fusion, memory coalescing and avoiding warp divergence

  • Experience writing quantised and mixed-precision kernels

  • Solid grasp of parallel algorithm design and numerical precision tradeoffs

  • Bonus: transformer/attention-style kernels or autoregressive decoding; building high-performance GPU libraries from scratch; HPC or latency-critical performance engineering; multi-GPU or multi-node kernel-level optimisation; comfortable reading PTX/SASS to validate kernel efficiency

  • Competitive salary and eligibility for equity

  • Biannual bonus scheme

  • Fully expensed tech to match your needs

  • Private health insurance

  • Breakfast and dinner allowance for office-based employees

As we hire globally, benefits vary by location.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior CUDA Engineer — Equity + Biannual Bonus
Senior CUDA Engineer — Equity + Biannual Bonus

Fuse Energy • United States

Remote
USD 140,000 - 210,000
Biannual bonus scheme
Fully expensed tech
Private health insurance
+1
AI Inference Engineer
AI Inference Engineer

Fuse Energy • United States

Remote
USD 200,000 - 350,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Software Engineer - GPU Kernel
Software Engineer - GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
Senior Software Engineer, C++ and CUDA - Analytics and Data Intelligence
Senior Software Engineer, C++ and CUDA - Analytics and Data Intelligence

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Software Engineer - GPU Kernels
Software Engineer - GPU Kernels

Baseten • United States

Remote
USD 140,000 - 210,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
Software Engineer - GPU Kernels
Software Engineer - GPU Kernels

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Software Engineer - GPU Kernels
Software Engineer - GPU Kernels

Baseten • New York (NY)

On-site
USD 180,000 - 360,000
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Paid parental leave
+2