GPU Kernel Engineer — Fast ML Training & Inference

Tilde Research

Palo Alto (CA)

On-site

USD 150,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce latency across models and infrastructure.

This role emphasizes performance-aware software, scalable kernels, and hands-on experimentation with CUDA, PyTorch, and Triton.

Qualifications

  • Experience in deep learning or related research areas.
  • Exceptional capability in ML kernel work (open source contributions, blogs, or similar).
  • Familiarity with PyTorch, Triton/TK/TileLang (any 2+), CUDA, and GPU architecture.

Responsibilities

  • Design, implement, and optimize GPU kernels for core model operations.
  • Collaborate with ML researchers to prototype and scale model architectures.
  • Contribute to system-wide efforts to improve efficiency and throughput beyond kernel-level work.

Skills

Deep learning
ML kernel development
PyTorch
CUDA familiarity
GPU architecture knowledge
Strong written communication
Quick learner
Problem solving

Tools

CUDA toolkit
Triton
TileLang
TK
Open source tooling

Job description

Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce latency across models and infrastructure.

This role emphasizes performance-aware software, scalable kernels, and hands-on experimentation with CUDA, PyTorch, and Triton.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kernel Engineer (Internship and Full-time)
Kernel Engineer (Internship and Full-time)

Tilde Research • Palo Alto (CA)

On-site
USD 150,000 - 260,000
GPU Kernel Engineer — Fast ML Training
GPU Kernel Engineer — Fast ML Training

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Kernel Architect for High-Performance AI Inference
GPU Kernel Architect for High-Performance AI Inference

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 120,000 - 180,000
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
KERNEL ENGINEER
KERNEL ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Kernel Engineer — High-Performance ML at Scale
GPU Kernel Engineer — High-Performance ML at Scale

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% medical, dental, and vision insurance coverage
Flexible PTO policy including a Winter Break
+2
ML Systems Engineer: Optimizing Training & GPU Kernels
ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
GPU Kernel Engineer: Build Fast AI Inference at Scale
GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
GPU Kernel Engineer for Distributed ML & Inference
GPU Kernel Engineer for Distributed ML & Inference

Advanced Micro Devices • Bellevue (WA), Northern (KY)

Hybrid
USD 140,000 - 190,000