Senior CUDA Engineer — Equity + Biannual Bonus

Fuse Energy

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Biannual bonus scheme
Fully expensed tech
Private health insurance
Breakfast and dinner allowance

Job summary

Fuse Energy is hiring a CUDA Engineer to write and optimise low-level GPU code for transformer inference workloads. You will design custom CUDA kernels, tune memory and compute bottlenecks, and push throughput across our fleet.

You have 4+ years in production CUDA, deep GPU microarchitecture knowledge, and experience with kernel fusion, mixed precision, and autoregressive decoding. We offer competitive salary, equity, and private health benefits.

Qualifications

  • 4+ years writing production CUDA code with performance kernels.
  • Deep understanding of GPU microarchitecture: warps, occupancy, registers.
  • Strong CUDA C++ skills with streams and asynchronous execution.
  • Hands-on profiling to diagnose compute-bound vs memory-bound bottlenecks.
  • Experience with kernel fusion, memory coalescing, warp-divergence avoidance.
  • Experience writing quantised and mixed-precision kernels.
  • Knowledge of parallel algorithms and numerical precision tradeoffs.
  • Bonus: transformer/attention-style kernels or autoregressive decoding; multi-GPU or multi-node kernel-level optimisation.

Responsibilities

  • Write and optimise custom CUDA kernels for core transformer inference operations
  • Profile kernels to identify bottlenecks in occupancy, memory throughput, and warp divergence
  • Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines
  • Optimise memory access patterns and manage the memory hierarchy for maximum bandwidth utilisation
  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprint
  • Build and tune caching mechanisms for efficient autoregressive decoding
  • Tune kernel launch configurations for target GPU architectures
  • Benchmark kernels against baselines and drive throughput and latency improvements
  • Write tests for CUDA code to catch performance and correctness regressions
  • Maintain internal CUDA libraries and contribute to coding standards and docs
  • 4+ years writing production CUDA code, with track record of shipping performance-critical kernels
  • Deep understanding of GPU microarchitecture: warps, occupancy, register pressure and memory hierarchy
  • Strong CUDA C++ skills, including streams and asynchronous execution
  • Hands-on experience profiling to diagnose compute-bound vs memory-bound bottlenecks
  • Experience with kernel fusion, memory coalescing and avoiding warp divergence
  • Experience writing quantised and mixed-precision kernels
  • Solid grasp of parallel algorithm design and numerical precision tradeoffs
  • Bonus: transformer/attention-style kernels or autoregressive decoding; building high-performance GPU libraries from scratch; HPC or latency-critical performance engineering; multi-GPU or multi-node kernel-level optimisation

Skills

CUDA C++
Kernel profiling
Kernel fusion
Memory optimization
Quantised kernels
Mixed precision
Autoregressive decoding
Multi-GPU kernels
Latency optimization

Job description

Fuse Energy is hiring a CUDA Engineer to write and optimise low-level GPU code for transformer inference workloads. You will design custom CUDA kernels, tune memory and compute bottlenecks, and push throughput across our fleet.

You have 4+ years in production CUDA, deep GPU microarchitecture knowledge, and experience with kernel fusion, mixed precision, and autoregressive decoding. We offer competitive salary, equity, and private health benefits.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

CUDA Engineer
CUDA Engineer

Fuse Energy • United States

Remote
USD 140,000 - 210,000
Biannual bonus scheme
Fully expensed tech
Private health insurance
+1
AI Inference Engineer
AI Inference Engineer

Fuse Energy • United States

Remote
USD 200,000 - 350,000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1
Software Engineer - GPU Kernels
Software Engineer - GPU Kernels

Baseten • United States

Remote
USD 140,000 - 210,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Senior Software Engineer - CUDA Driver
Senior Software Engineer - CUDA Driver

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Equity
Benefits package
Senior Software Engineer, CUDA UMD - Graphs and GPU Sharing
Senior Software Engineer, CUDA UMD - Graphs and GPU Sharing

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

On-site
USD 184,000 - 357,000
Senior Software Engineer, C++ and CUDA - Analytics and Data Intelligence
Senior Software Engineer, C++ and CUDA - Analytics and Data Intelligence

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Remote CUDA GPU Engineer for High-Performance AI/HPC
Remote CUDA GPU Engineer for High-Performance AI/HPC

Bright Vision Technologies • Woodbridge Township (NJ)

Remote
USD 100,000 - 175,000
Principal Engineer, CUDA UMD - GPU Kernel Scheduling
Principal Engineer, CUDA UMD - GPU Kernel Scheduling

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Benefits