Research Engineer - CUDA Kernel Engineering

Voltai

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Voltai is building world models and AI-enabled semiconductor tools. You will develop, integrate, and optimize CUDA kernels to accelerate design and verification tasks across thousands of GPUs.

You will create tooling, benchmarks, and integration layers, collaborating with researchers to advance AI-driven hardware design and release kernels to open-source ecosystems. Youll work closely with researchers and engineers to push the limits of GPU utilization for compute-intensive workloads and

Qualifications

  • Writing and optimizing CUDA kernels for large-scale AI workloads.
  • Profiling and optimizing GPU performance for compute-heavy tasks.
  • Integrating custom kernels into PyTorch, Megatron, vLLM, TorchTitan.
  • Working with NVIDIA hardware and software stacks (Hopper/Blackwell, NVLink, NCCL, Triton).
  • Building GPU-accelerated primitives for graph reasoning and hardware sim.
  • Collaborating with AI researchers and semiconductor experts to translate workloads into high-performance GPU code.

Responsibilities

  • Develop, integrate, and optimize state-of-the-art CUDA kernels for AI models.
  • Power large-scale model training, inference, and reinforcement learning workloads.
  • Build tools, benchmarks, and integration layers to maximize GPU utilization.
  • Collaborate with researchers and engineers to push Voltai's AI+semiconductor objectives.

Skills

CUDA kernels
GPU profiling
Framework integration
NVIDIA hardware
GPU primitives
Research collaboration

Tools

PyTorch
Megatron
vLLM
TorchTitan
NCCL/Triton

Job description

About Voltai

Voltai is developing world models, and agents to learn, evaluate, plan, experiment, and interact with the physical world. We are starting out with understanding and building hardware; electronics systems and semiconductors where AI can design and create beyond human cognitive limits.

About the Team

Backed by Silicon Valley’s top investors, Stanford University, and CEOs/Presidents of Google, AMD, Broadcom, Marvell, etc. We are a team of previous Stanford professors, SAIL researchers, Olympiad medalists (IPhO, IOI, etc.), CTOs of Synopsys & GlobalFoundries, Head of Sales & CRO of Cadence, former US Secretary of Defense, National Security Advisor, and Senior Foreign‑Policy Advisor to four US presidents.

About the Role

You will develop, integrate, and optimize state‑of‑the‑art CUDA kernels to power AI models that accelerate semiconductor design and verification. Your work will enable large‑scale model training, inference, and reinforcement learning systems that reason about circuit layouts, generate and validate RTL, and optimize chip architectures — running efficiently across thousands of GPUs. You’ll build tools, performance benchmarks, and integration layers that push the limits of GPU utilization for compute‑intensive workloads in AI‑driven hardware design. Working closely with researchers and engineers, you’ll help make Voltai the world’s leading AI + semiconductor research organization. You’ll also release your kernels and tooling as contributions to the open‑source AI and HPC ecosystems.

You might thrive in this role if you have experience with
  • Writing and optimizing CUDA kernels for large‑scale AI workloads (attention, routing, graph‑based operations, physics‑inspired operators, etc.)
  • Profiling and optimizing GPU performance for custom compute or memory‑bound workloads
  • Integrating custom kernels into cutting‑edge training and inference frameworks (e.g., PyTorch, Megatron, vLLM, TorchTitan)
  • Working with the latest NVIDIA hardware and software stacks (Hopper, Blackwell, NVLink, NCCL, Triton)
  • Building GPU‑accelerated primitives for graph reasoning, symbolic computation, or hardware simulation tasks
  • Collaborating with AI researchers and semiconductor experts to translate domain‑specific workloads into high‑performance GPU code
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

CUDA Kernel Engineer for AI Hardware & Semiconductors
CUDA Kernel Engineer for AI Hardware & Semiconductors

Voltai • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Engineer - Post-Training
Research Engineer - Post-Training

Voltai • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Server Hardware Engineer
Server Hardware Engineer

Voltai • Palo Alto (CA)

On-site
USD 140,000 - 190,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Research Engineer - Mid-Training
Research Engineer - Mid-Training

Voltai • Palo Alto (CA)

On-site
USD 180,000 - 240,000
KERNEL ENGINEER
KERNEL ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
Machine Learning Engineer
Machine Learning Engineer

Voltai • Palo Alto (CA)

On-site
USD 180,000 - 240,000