Senior AI Triton & Distributed GPU Software Engineer

Advanced Micro Devices

San Jose (CA)

On-site

USD 180,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD is advancing Triton on AMD Instinct GPUs, delivering native distributed execution and communication to scale AI workloads across multiple GPUs. The role spans compiler, runtime, and hardware interactions to maximize performance and throughput in data centers and HPC environments.

You will collaborate with architecture, software, and performance teams to push the limits of Triton-based distributed AI on AMD hardware, contributing to open-source ecosystems and ROCm integration while guiding

Qualifications

  • Must have hands-on experience in GPU software, compilers, or distributed systems.
  • Strong knowledge of modern GPU architectures, memory hierarchy, interconnects.
  • Experience with distributed GPU libraries and multi-GPU programming.

Responsibilities

  • Design and implement native distributed communication and execution for Triton AMDGPU backend.
  • Develop compiler/runtime mechanisms for GPU-initiated communication and distributed execution.
  • Improve inter-GPU data movement, overlap of communication and computation, and memory hierarchy usage.
  • Optimize distributed Triton kernels and execution models for scalability on AMD Instinct GPUs.
  • Analyze, profile, and debug cross-stack issues across Triton, ROCm, and hardware.

Skills

Compiler development
Triton compiler
GPU architecture
GPU runtime
NCCL/RCCL/MPI
Multi-GPU scaling
CUDA/Triton
MLIR/LLVM
Profiling/Debugging
HIP/CUDA/ROCm
AI/HPC workloads
Open-source
Leadership/Communication

Education

Bachelor’s or Master’s Degree in Computer Engineering/CS/EE

Job description

AMD is advancing Triton on AMD Instinct GPUs, delivering native distributed execution and communication to scale AI workloads across multiple GPUs. The role spans compiler, runtime, and hardware interactions to maximize performance and throughput in data centers and HPC environments.

You will collaborate with architecture, software, and performance teams to push the limits of Triton-based distributed AI on AMD hardware, contributing to open-source ecosystems and ROCm integration while guiding

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Distributed AI & Triton on GPUs
Senior Software Engineer, Distributed AI & Triton on GPUs

Socket.dev • San Jose (CA)

On-site
USD 180,000 - 240,000
AMD benefits
Senior AI Triton & Distributed GPU Engineer
Senior AI Triton & Distributed GPU Engineer

AMD • San Jose (CA)

On-site
USD 180,000 - 240,000
Benefits at a glance
Senior AI GPU & Distributed-Systems Engineer
Senior AI GPU & Distributed-Systems Engineer

Advanced Micro Devices, Inc. (AMD) • United States

On-site
USD 150,000 - 190,000
Senior Software Engineer - AI
Senior Software Engineer - AI

Advanced Micro Devices, Inc. (AMD) • United States

On-site
USD 150,000 - 190,000
Senior Triton GPU Kernel Engineer for AI Workloads
Senior Triton GPU Kernel Engineer for AI Workloads

Socket.dev • San Jose (CA)

Hybrid
USD 180,000 - 300,000
AMD benefits at a glance
Senior GPU Kernel Engineer - Hybrid Triton AI
Senior GPU Kernel Engineer - Hybrid Triton AI

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 180,000 - 280,000
Sr. Software Engineer - AI Triton Communication
Sr. Software Engineer - AI Triton Communication

AMD • San Jose (CA)

On-site
USD 180,000 - 240,000
Benefits at a glance
Sr. Software Engineer - AI Triton Communication
Sr. Software Engineer - AI Triton Communication

Socket.dev • San Jose (CA)

On-site
USD 180,000 - 240,000
AMD benefits
Sr. Software Engineer - AI Triton Communication
Sr. Software Engineer - AI Triton Communication

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 300,000
Senior GPU Kernel Engineer – Triton, Hybrid
Senior GPU Kernel Engineer – Triton, Hybrid

AMD • United States

Hybrid
USD 180,000 - 260,000
Benefits at a glance
Hybrid work model