CUDA Engineer: High-Perf GPU Kernels for Inference

Fuse Energy

United States

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Biannual bonus
Equity eligibility
Fully expensed tech
Private health insurance
Meals allowance

Job summary

Fuse Energy seeks a CUDA Engineer to write and optimise low-level GPU code powering our transformer inference workloads. You will design custom CUDA kernels, tune memory patterns, and push maximum throughput across our GPU fleet.

Responsibilities include kernel fusion, mixed-precision arithmetic, and efficient autoregressive decoding. You will profile, test, and maintain CUDA libraries while collaborating with a fast-moving AI-enabled energy team.

Qualifications

  • 4+ years writing production CUDA code.
  • Deep understanding of GPU microarchitecture: warps, occupancy, registers and memory hierarchy.
  • Strong CUDA C++ skills, including streams and asynchronous execution.
  • Experience with kernel fusion, memory coalescing and avoiding warp divergence.
  • Experience writing quantised and mixed-precision kernels.

Responsibilities

  • Write and optimise custom CUDA kernels for core transformer inference operations.
  • Profile kernels to identify bottlenecks in occupancy, memory throughput and warp divergence.
  • Apply kernel fusion to reduce memory round-trips and launch overhead.
  • Optimise memory access patterns and manage memory hierarchy for max bandwidth.
  • Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency.
  • Build and tune caching mechanisms for autoregressive decoding.
  • Tune kernel launch configurations for target GPU architectures.
  • Benchmark kernels against baselines and drive throughput improvements.
  • Write tests for CUDA code to catch performance and correctness regressions.

Skills

CUDA C++
GPU profiling
Kernel optimization
Quantisation & mixed precision
Parallel algorithm design
HPC latency engineering
Multi-GPU / multi-node

Job description

Fuse Energy seeks a CUDA Engineer to write and optimise low-level GPU code powering our transformer inference workloads. You will design custom CUDA kernels, tune memory patterns, and push maximum throughput across our GPU fleet.

Responsibilities include kernel fusion, mixed-precision arithmetic, and efficient autoregressive decoding. You will profile, test, and maintain CUDA libraries while collaborating with a fast-moving AI-enabled energy team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Engineer
CUDA Engineer

Fuse Energy • United States

On-site
USD 140,000 - 210,000
Biannual bonus
Equity eligibility
Fully expensed tech
+2
Inference Runtime Performance Engineer — GPU Kernels
Inference Runtime Performance Engineer — GPU Kernels

iFrame Corporation • San Francisco (CA)

Remote
USD 220,000 - 360,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Senior CUDA Kernel Engineer for High-Performance Inference
Senior CUDA Kernel Engineer for High-Performance Inference

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+3
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Founding AI Inference Engineer
Founding AI Inference Engineer

Fuse Energy • United States

On-site
USD 180,000 - 300,000
Competitive salary and equity
Biannual bonus
Fully expensed tech to match needs
+2
AI Inference Engineer
AI Inference Engineer

Fuse Energy • United States

On-site
USD 180,000 - 300,000
Competitive salary and equity
Biannual bonus
Fully expensed tech to match needs
+2
Inference Performance Engineer — GPU Kernels & Systems
Inference Performance Engineer — GPU Kernels & Systems

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Distinguished Inference Engineer
Distinguished Inference Engineer

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 320,000