Senior GPU Kernel Engineer - Hybrid Triton AI

Advanced Micro Devices

San Jose (CA)

Hybrid

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Advanced Micro Devices seeks a highly skilled engineer to advance Triton-based GPU kernels for AI workloads. You will work across Triton/Gluon kernels, PyTorch, vLLM, and ROCm stacks to push performance on AMD Instinct GPUs.

You will partner with researchers, compiler engineers, and hardware teams to design and optimize cutting-edge kernels, improving throughput and efficiency for large-scale AI models while contributing to open-source ecosystems.

Qualifications

  • Deep experience in GPU kernel development and AI/ML workloads.
  • Hands-on Triton kernels and GPU architectures.
  • Strong collaboration across research, compiler, and hardware teams.

Responsibilities

  • Design, research, implement, and optimize high-performance matmul, attention (flash, paged, grouped-query), MoE, and fully fused transformer kernels using Triton, targeting large-scale LLM and multimodal workloads.
  • Own and productionize critical Triton/Gluon kernels within vLLM and SGL (e.g., paged attention, extend attention, MoE, quantized kernels).
  • Partner closely with compiler engineers to develop and maintain the Triton AMD backend across ROCm and the LLVM AMDGPU stack, targeting CDNA and next-generation architectures.
  • Drive deep kernel-level optimizations across the AMD memory hierarchy (LDS, L2, HBM), wavefront execution (wave32/wave64), vectorization, MFMA utilization, occupancy tuning, and instruction scheduling.
  • Perform rigorous profiling and microbenchmarking led optimization on AMD Instinct GPUs using hardware counters and tracing tools; root-cause bottlenecks in memory bandwidth, latency hiding, synchronization, and register pressure.
  • Debug and resolve performance and correctness issues end-to-end across PyTorch, vLLM/SGL runtimes, Triton IR/MLIR, ROCm runtime, and the LLVM AMDGPU backend.
  • Contribute to open-source Triton, LLVM, and ROCm ecosystems

Skills

GPU kernel development
Compiler backends
Performance engineering
AI/ML workloads

Education

Bachelor's or Master's degree in CS/CE/EE

Tools

Triton
ROCm
LLVM
SGLang
Torch
MLIR

Job description

Advanced Micro Devices seeks a highly skilled engineer to advance Triton-based GPU kernels for AI workloads. You will work across Triton/Gluon kernels, PyTorch, vLLM, and ROCm stacks to push performance on AMD Instinct GPUs.

You will partner with researchers, compiler engineers, and hardware teams to design and optimize cutting-edge kernels, improving throughput and efficiency for large-scale AI models while contributing to open-source ecosystems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Triton GPU Kernel Engineer for AI Workloads
Senior Triton GPU Kernel Engineer for AI Workloads

Socket.dev • San Jose (CA)

Hybrid
USD 180,000 - 300,000
AMD benefits at a glance
Senior GPU Kernel Engineer – Triton, Hybrid
Senior GPU Kernel Engineer – Triton, Hybrid

AMD • United States

Hybrid
USD 180,000 - 260,000
Benefits at a glance
Hybrid work model
Senior AI Triton & Distributed GPU Engineer
Senior AI Triton & Distributed GPU Engineer

AMD • San Jose (CA)

On-site
USD 180,000 - 240,000
Benefits at a glance
Senior Software Engineer, Distributed AI & Triton on GPUs
Senior Software Engineer, Distributed AI & Triton on GPUs

Socket.dev • San Jose (CA)

On-site
USD 180,000 - 240,000
AMD benefits
Senior AI Triton & Distributed GPU Software Engineer
Senior AI Triton & Distributed GPU Software Engineer

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 300,000
Sr. Software Engineer - AI Triton Kernels
Sr. Software Engineer - AI Triton Kernels

Socket.dev • San Jose (CA)

Hybrid
USD 180,000 - 300,000
AMD benefits at a glance
Sr. Software Engineer - AI Triton Kernels
Sr. Software Engineer - AI Triton Kernels

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 180,000 - 280,000
Sr. Software Engineer - AI Triton Kernels
Sr. Software Engineer - AI Triton Kernels

AMD • United States

Hybrid
USD 180,000 - 260,000
Benefits at a glance
Hybrid work model
Sr. Software Engineer - AI Triton Kernels
Sr. Software Engineer - AI Triton Kernels

AMD • San Jose (CA)

On-site
USD 190,000 - 270,000
Senior GPU/AI Systems Engineer - Performance & ML
Senior GPU/AI Systems Engineer - Performance & ML

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits