CUDA Engineer

Fuse Energy

United States

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Biannual bonus
Equity eligibility
Fully expensed tech
Private health insurance
Meals allowance

Job summary

Fuse Energy seeks a CUDA Engineer to write and optimise low-level GPU code powering our transformer inference workloads. You will design custom CUDA kernels, tune memory patterns, and push maximum throughput across our GPU fleet.

Responsibilities include kernel fusion, mixed-precision arithmetic, and efficient autoregressive decoding. You will profile, test, and maintain CUDA libraries while collaborating with a fast-moving AI-enabled energy team.

Qualifications

  • 4+ years writing production CUDA code.
  • Deep understanding of GPU microarchitecture: warps, occupancy, registers and memory hierarchy.
  • Strong CUDA C++ skills, including streams and asynchronous execution.
  • Experience with kernel fusion, memory coalescing and avoiding warp divergence.
  • Experience writing quantised and mixed-precision kernels.

Responsibilities

  • Write and optimise custom CUDA kernels for core transformer inference operations.
  • Profile kernels to identify bottlenecks in occupancy, memory throughput and warp divergence.
  • Apply kernel fusion to reduce memory round-trips and launch overhead.
  • Optimise memory access patterns and manage memory hierarchy for max bandwidth.
  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency.
  • Build and tune caching mechanisms for autoregressive decoding.
  • Tune kernel launch configurations for target GPU architectures.
  • Benchmark kernels against baselines and drive throughput improvements.
  • Write tests for CUDA code to catch performance and correctness regressions.

Skills

CUDA C++
GPU profiling
Kernel optimization
Quantisation & mixed precision
Parallel algorithm design
HPC latency engineering
Multi-GPU / multi-node

Job description

Fuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.

We've raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.

We're building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're looking for a CUDA Engineer to write and optimise the low-level GPU code that powers our inference workloads: designing custom CUDA kernels, tuning performance across memory bandwidth and compute bottlenecks, and squeezing maximum throughput out of every GPU in our fleet, working at the level of SMs, warps and memory hierarchies.

Responsibilities
  • Write and optimise custom CUDA kernels for core transformer inference operations
  • Profile kernels to identify and eliminate bottlenecks in occupancy, memory throughput and warp divergence
  • Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines
  • Optimise memory access patterns and manage the memory hierarchy for maximum bandwidth utilisation
  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprint
  • Build and tune caching mechanisms for efficient autoregressive decoding
  • Tune kernel launch configurations for target GPU architectures
  • Benchmark kernels against existing baselines and drive measurable throughput and latency improvements
  • Write tests for CUDA code to catch performance and correctness regressions
  • Maintain internal CUDA libraries and contribute to team coding standards and documentation
  • 4+ years writing production CUDA code, with a track record of shipping performance-critical kernels
  • Deep understanding of GPU microarchitecture: warps, occupancy, register pressure and memory hierarchy
  • Strong CUDA C++ skills, including streams and asynchronous execution
  • Hands‑on experience profiling to diagnose compute-bound vs memory-bound bottlenecks
  • Experience with kernel fusion, memory coalescing and avoiding warp divergence
  • Experience writing quantised and mixed-precision kernels
  • Solid grasp of parallel algorithm design and numerical precision tradeoffs
  • Bonus: transformer/attention-style kernels or autoregressive decoding; building high-performance GPU libraries from scratch; HPC or latency‑critical performance engineering; multi‑GPU or multi‑node kernel‑level optimisation; comfortable reading PTX/SASS to validate kernel efficiency
  • Competitive salary and eligibility for equity
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Private health insurance
  • Breakfast and dinner allowance for office‑based employees

As we hire globally, benefits vary by location.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Engineer: High-Perf GPU Kernels for Inference
CUDA Engineer: High-Perf GPU Kernels for Inference

Fuse Energy • United States

On-site
USD 140,000 - 210,000
Biannual bonus
Equity eligibility
Fully expensed tech
+2
AI Inference Engineer
AI Inference Engineer

Fuse Energy • United States

On-site
USD 180,000 - 300,000
Competitive salary and equity
Biannual bonus
Fully expensed tech to match needs
+2
Senior Software Engineer, C++ and CUDA - Analytics and Data Intelligence
Senior Software Engineer, C++ and CUDA - Analytics and Data Intelligence

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Senior Software Engineer, CUDA UMD - Graphs and GPU Sharing
Senior Software Engineer, CUDA UMD - Graphs and GPU Sharing

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Senior Software Engineer CUDA UMD - GPU Kernel Scheduling
Senior Software Engineer CUDA UMD - GPU Kernel Scheduling

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Software Engineer CUDA UMD - GPU Kernel Scheduling
Senior Software Engineer CUDA UMD - GPU Kernel Scheduling

NVIDIA • Jasper (AL)

On-site
USD 152,000 - 288,000
Equity
Benefits
Member of Technical Staff - GPU Performance Engineer San Francisco · Remote · Boston · Hybrid
Member of Technical Staff - GPU Performance Engineer San Francisco · Remote · Boston · Hybrid

Liquid AI, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity in unicorn-stage company
Health premiums paid
401(k) matching up to 4%
+2
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1