CUDA Kernel Engineer (Remote US)

Pragmatike

New Jersey

On-site

USD 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary & equity options
Sign-on bonus
Health, Dental, and Vision
401k

Job summary

Pragmatike is hiring a CUDA Kernel Engineer to develop and optimize NVIDIA CUDA kernels for a leading AI startup. This remote role focuses on maximizing GPU performance and throughput for high-scale AI systems. Candidates should have substantial experience with CUDA and a profound understanding of GPU architecture. The position includes competitive salary, equity, sign-on bonuses, and comprehensive health benefits. Join a fast-paced environment that’s shaping the future of AI.

Qualifications

  • Proven track record building NVIDIA CUDA kernels from scratch.
  • Strong ability to optimize kernels with advanced techniques.
  • Deep understanding of GPU memory hierarchy and performance bottlenecks.

Responsibilities

  • Design, implement, and optimize custom CUDA kernels for NVIDIA GPUs.
  • Profile GPU workloads using advanced profiling tools.
  • Collaborate with AI systems and model acceleration teams.

Skills

NVIDIA CUDA kernels optimization
GPU performance analysis
C++ proficiency
Memory management techniques

Tools

Nsight Compute
nvprof
CUDA-MEMCHECK

Job description

About The Role

Pragmatike is hiring on behalf of a fast‑growing AI startup recognized as a Top 10 GenAI company by GTM Capital, founded by MIT CSAIL researchers.

Location: Remote US
Start date: ASAP
Languages: English (required)

We are searching for a CUDA Kernel Engineer who has hands‑on experience developing and optimizing NVIDIA CUDA kernels from scratch. You will work on the GPU performance layer powering large-scale, high-throughput AI systems used by Fortune 500 customers. This role is ideal for someone who deeply understands NVIDIA GPU architecture, memory hierarchy, warp‑level execution, and profiling workflows – not someone coming from generic hardware, FPGA, or non‑NVIDIA compute backgrounds. You will directly influence the GPU efficiency, throughput, and scalability of mission‑critical AI systems.

What You’ll Do
  • Design, implement, and optimize custom CUDA kernels for NVIDIA GPUs, with a focus on maximizing occupancy, memory throughput, and warp efficiency.
  • Profile GPU workloads using tools such as Nsight Compute, Nsight Systems, nvprof, and CUDA‑MEMCHECK.
  • Analyze and eliminate performance bottlenecks including warp divergence, uncoalesced memory access, register pressure, and PCIe transfer overhead.
  • Improve GPU memory pipelines (global, shared, L2, texture memory) and ensure proper memory coalescing.
  • Collaborate closely with AI systems, model acceleration, and backend distributed systems teams.
  • Contribute to GPU architecture decisions, kernel libraries, and internal performance‑engineering best practices.
What We’re Looking For
  • Proven track record building NVIDIA CUDA kernels from scratch, not just calling existing libraries.
  • Strong ability to optimize kernels (tiling strategies, occupancy tuning, shared memory design, warp scheduling).
  • Deep understanding of CUDA threads, warps, blocks, and grids, GPU memory hierarchy and memory coalescing, as well as warp divergence (how to detect, analyze, and mitigate it).
  • Experience diagnosing PCIe bottlenecks and optimizing host‑device transfers (pinned memory, streams, batching, overlap).
  • Familiarity with C++, CUDA runtime APIs, and GPU debugging/profiling tooling.
Bonus Points
  • Experience with multi‑GPU or distributed GPU systems (NCCL, NVLink, MIG).
  • Background in GPU acceleration for ML frameworks or HPC workloads.
  • Knowledge of model inference optimization (TensorRT, CUDA Graphs, CUTLASS).
  • Exposure to compiler‑level optimization or PTX/SASS analysis.
  • Startup experience or comfort working in fast‑moving, ambiguous environments.
Benefits
  • Competitive salary & equity options
  • Sign‑on bonus
  • Health, Dental, and Vision
  • 401k
EEO Statement

Pragmatike is an Equal Opportunity Employer and is committed to providing equal employment opportunities to all applicants without discrimination. We recruit on behalf of our clients and prohibit discrimination and harassment based on race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws. This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training. We are committed to a fair and inclusive hiring process. We process your personal data solely for recruitment purposes, in accordance with applicable privacy laws, and maintain reasonable safeguards to protect your information. Your data may be shared with our client(s) for hiring consideration, but will not be disclosed to third parties outside of the recruitment process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote CUDA Kernel Engineer – Optimize GPU Performance
Remote CUDA Kernel Engineer – Optimize GPU Performance

Pragmatike • Cambridge (MA)

On-site
USD 150,000 - 230,000
Salary + equity
Sign-on bonus
Health/Dental/Vision
+1
Remote CUDA Kernel Engineer - Optimize AI GPU Pipelines
Remote CUDA Kernel Engineer - Optimize AI GPU Pipelines

Pragmatike • Town of Florida (NY)

On-site
USD 120,000 - 150,000
Competitive salary & equity options
Sign-on bonus
Health, Dental, and Vision
+1
Remote CUDA Kernel Engineer - AI GPU Performance
Remote CUDA Kernel Engineer - AI GPU Performance

Pragmatike • New Jersey

On-site
USD 110,000 - 150,000
Founding GPU Kernel Engineer
Founding GPU Kernel Engineer

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
GPU Performance / Kernel Engineer
GPU Performance / Kernel Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan
Annual bonus
+1
37303 HD - CUDA Engineering Expert
37303 HD - CUDA Engineering Expert

Cephas Consultancy Services Private Limited • California (MO)

Hybrid
USD 120,000 - 180,000
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

NVIDIA • Town of Texas (WI)

On-site
USD 124,000 - 196,000
Equity
Benefits
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

NVIDIA • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity
Benefits