Sr. Software Engineer - AI Triton Communication

AMD

United States

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

AMD is seeking a seasoned engineer to advance Triton distributed execution and communication on AMD Instinct GPUs. You will work across compiler, runtime, and hardware layers to enable scalable multi-GPU training and inference, with close collaboration across architecture, software, and performance teams.

You will design native distributed capabilities, optimize inter-GPU data movement, and contribute to open-source Triton and ROCm ecosystems.

Qualifications

  • Experience in compiler development and GPU software.
  • Proven able to optimize workloads across multi-GPU systems.
  • Strong understanding of GPU memory hierarchy and interconnects.
  • Familiarity with distributed GPU communication libraries.

Responsibilities

  • Design and develop native distributed communication and execution within Triton AMDGPU backend.
  • Implement compiler/runtime mechanisms for GPU-initiated communication and collective ops.
  • Optimize inter-GPU data movement and scheduling for AI workloads.
  • Collaborate with architecture, compiler, runtime and performance teams.

Skills

Compiler development
GPU software
Distributed systems
Performance engineering
Triton
Runtime
GPU architectures
Inter-GPU interconnects

Tools

Triton compiler
ROCm
NCCL

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we're looking for talent who feel the same: people who want to leave the planet better than they found it, those who don't shy away from humanity's challenges but are determined to help solve them.

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you're designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger - technology that moves the world forward.

THE ROLE:

Triton is a widely adopted language and compiler for high-performance GPU kernels, powering major AI frameworks such as PyTorch, vLLM, and SGLang. As AI workloads increasingly scale across multiple GPUs and nodes, first-class support for distributed execution and communication in Triton is strategically critical to enabling efficient large-scale training and inference on AMD Instinct Accelerators. AMD GPUs are an official Triton backend, and delivering industry-leading distributed performance and scalability on AMD Instinct accelerators is a key priority. The performance, scalability, and usability of Triton directly impact the competitiveness of AMD hardware in large-scale AI deployments. In this role, you will advance the Triton compiler and runtime stack for AMD CDNA and next-generation GPU architectures by building native distributed execution and communication capabilities. You will develop compiler and runtime infrastructure that enables efficient inter-GPU communication, scalable execution, and optimal hardware utilization. You will work across compiler, runtime, and hardware layers, collaborating closely with GPU architecture and software teams to help establish AMD GPUs as a best-in-class platform for Triton-based distributed AI.

THE PERSON:

The ideal candidate has deep expertise in GPU architecture, compiler technologies, and distributed GPU systems, with proven experience optimizing workloads at multi-GPU scale. You are comfortable working across the full execution stack - from compiler and runtime to hardware - and understand how GPU execution, memory hierarchy, and inter-GPU communication impact performance. You have experience working close to the GPU runtime, communication stack, or compiler backend, and are motivated to build native distributed execution and communication capabilities tightly integrated with the compiler and runtime to maximize scalability and hardware utilization. You thrive on solving complex system-level performance challenges and delivering scalable, high-performance GPU infrastructure.

KEY RESPONSIBILITIES:
  • Design and develop native distributed communication and execution capabilities within the Triton AMDGPU backend, enabling scalable multi-GPU execution for large-scale AI workloads
  • Design and implement Triton compiler and runtime mechanisms for native GPU-initiated communication, including collective operations, remote memory access, synchronization and distributed execution primitives
  • Drive performance optimization across compute and communication, including inter-GPU data movement, communication/computation overlap, memory hierarchy utilization, and GPU-driven scheduling efficiency
  • Develop and optimize distributed Triton kernels and execution models to achieve high performance, scalability, and efficient hardware utilization for AI workloads
  • Analyze, profile and debug complex cross-stack issues spanning Triton compiler, runtime, ROCm stack, and GPU hardware execution
  • Collaborate closely with GPU architecture, compiler, runtime, and performance teams to co-design and enable next-generation distributed GPU programming and execution capabilities
  • Contribute to open-source Triton and ROCm distributed ecosystem, driving innovation in distributed GPU computing
PREFERRED EXPERIENCE:
  • Deep experience in compiler development, GPU software, distributed systems, or performance engineering
  • Familiarity or hands-on experience with Triton compiler and runtime
  • Deep understanding of modern GPU architectures, including execution model, memory hierarchy (LDS, L2, HBM), scheduling, occupancy, and hardware performance characteristics
  • Good understanding of GPU runtime systems, communication stacks, and multi-GPU interconnects such as XGMI, NVLink, PCIe, or InfiniBand and their performance implications
  • Familiarity with distributed GPU communication libraries such as RCCL, NCCL, NVSHMEM, rocSHMEM, or MPI and similar technologies
  • Experience developing, optimizing, and scaling workloads across multiple GPUs, including inter-GPU communication, synchronization, and communication/computation overlap
  • Strong experience with GPU programming using Triton, HIP, CUDA, or similar parallel programming environments
  • Strong knowledge of MLIR and/or LLVM internals
  • Ex
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Software Engineer - AI Triton Communication
Sr. Software Engineer - AI Triton Communication

AMD • San Jose (CA)

On-site
USD 180,000 - 240,000
Benefits at a glance
Sr. Software Engineer - AI Triton Communication
Sr. Software Engineer - AI Triton Communication

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 300,000
Senior AI Triton & Distributed GPU Engineer
Senior AI Triton & Distributed GPU Engineer

AMD • San Jose (CA)

On-site
USD 180,000 - 240,000
Benefits at a glance
Senior AI Triton Distributed GPU Engineer
Senior AI Triton Distributed GPU Engineer

AMD • United States

On-site
USD 150,000 - 210,000
Senior AI Triton & Distributed GPU Software Engineer
Senior AI Triton & Distributed GPU Software Engineer

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 300,000
Sr. Software Engineer - AI Triton Kernels
Sr. Software Engineer - AI Triton Kernels

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 180,000 - 280,000
Sr. Software Engineer - AI Triton Kernels
Sr. Software Engineer - AI Triton Kernels

AMD • United States

Hybrid
USD 180,000 - 260,000
Benefits at a glance
Hybrid work model
Sr. Software Engineer - AI Triton Kernels
Sr. Software Engineer - AI Triton Kernels

AMD • San Jose (CA)

On-site
USD 190,000 - 270,000
Senior GPU Kernel Engineer - Hybrid Triton AI
Senior GPU Kernel Engineer - Hybrid Triton AI

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 180,000 - 280,000
Senior GPU Kernel Engineer – Triton, Hybrid
Senior GPU Kernel Engineer – Triton, Hybrid

AMD • United States

Hybrid
USD 180,000 - 260,000
Benefits at a glance
Hybrid work model