Kernel Engineer (Internship and Full-time)

Tilde Research

Palo Alto (CA)

On-site

USD 150,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce latency across models and infrastructure.

This role emphasizes performance-aware software, scalable kernels, and hands-on experimentation with CUDA, PyTorch, and Triton.

Qualifications

  • Experience in deep learning or related research areas.
  • Exceptional capability in ML kernel work (open source contributions, blogs, or similar).
  • Familiarity with PyTorch, Triton/TK/TileLang (any 2+), CUDA, and GPU architecture.

Responsibilities

  • Design, implement, and optimize GPU kernels for core model operations.
  • Collaborate with ML researchers to prototype and scale model architectures.
  • Contribute to system-wide efforts to improve efficiency and throughput beyond kernel-level work.

Skills

Deep learning
ML kernel development
PyTorch
CUDA familiarity
GPU architecture knowledge
Strong written communication
Quick learner
Problem solving

Tools

CUDA toolkit
Triton
TileLang
TK
Open source tooling

Job description

Tilde Research is a moonshot AI lab advancing mechanistic interpretability, new architectures, and pretraining science. We build foundational understanding of models to advance the frontier of intelligence.

About The Role

As a Kernel Engineer at Tilde, you'll design, implement, and optimize high-performance GPU kernels that are critical to scaling our training and inference workloads. Your work will enable faster iteration cycles, higher throughput, and lower latency. You'll work closely with ML researchers and engineers to co-design models and infrastructure that are deeply performance-aware, and help push the limits of what current hardware can support.

What You Might Work On
  • Design, develop, and tune custom GPU kernels for core model operations
  • Work with ML engineers to prototype and scale novel model architectures
  • Contribute to system-wide efforts to improve efficiency and throughput, beyond just kernel-level optimizations
You're a Good Fit If You
  • Have experience in deep learning or related research areas
  • Have demonstrated exceptional capability in working on ML kernels. This can include:
    • Strong open source contributions
    • Thoughtful technical blog posts/work logs
    • Previous experience working on hardware-aligned algorithms
  • Deep familiarity with PyTorch, Triton/TK/TileLang (>1 of), basic familiarity with CUDA, and knowledge of GPU architecture.
  • Communicate clearly and effectively, both verbally and in writing
  • Strong algorithmic thinker
  • Are able to learn quickly
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Kernel Engineer — Fast ML Training & Inference
GPU Kernel Engineer — Fast ML Training & Inference

Tilde Research • Palo Alto (CA)

On-site
USD 150,000 - 260,000
ML Engineer (Internship and Full-time)
ML Engineer (Internship and Full-time)

Tilde Research • Palo Alto (CA)

On-site
USD 140,000 - 190,000
KERNEL ENGINEER
KERNEL ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Engineer
ML Engineer

Tilde Research • San Francisco (CA)

On-site
USD 180,000 - 230,000
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
Research Engineer - CUDA Kernel Engineering
Research Engineer - CUDA Kernel Engineering

Voltai • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Engineer, Infrastructure, Kernels
Research Engineer, Infrastructure, Kernels

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3