Principal System Software Architect, AI/GPU Platforms

AMD

Austin (TX)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD is seeking a Systems Software Architect to join the MI400-class subsystems team, focusing on next‑generation, rack‑scale AI platforms. You will define software stack interfaces across kernel drivers, ROCm runtime, and firmware, collaborating with silicon teams to deliver scalable performance.

The role emphasizes reliability and performance at scale, with experience in GPU compute ecosystems and collaboration with AI workloads. Austin, TX and Santa Clara, CA locations apply.

Qualifications

  • Linux memory management and heterogeneous memory management knowledge.
  • GPU/DRM driver development experience.
  • Cache coherence and memory consistency knowledge.
  • GPU networking technologies incl. NVLink, UALink, RDMA.
  • Scale-up/scale-out networking and RCCL/NCCL experience.
  • Familiarity with data-center fabrics like Infinity Fabric/InfiniBand/RoCE.

Responsibilities

  • Own end-to-end system software architecture for MI400-class subsystems (GPU memory management, scheduling, RAS).
  • Drive architecture across kernel driver, user-runtime, firmware interfaces, ROCm platform.
  • Collaborate with silicon/SoC architects during pre-silicon definition to shape interfaces and contracts.
  • Define software strategy for multi-GPU and rack-scale topologies and interconnects.
  • Establish reliability, availability, serviceability, telemetry, and graceful degradation at scale.
  • Identify bottlenecks in launch path, memory subsystem, and communication paths and propose software solutions.
  • Produce architecture specifications, reference designs, and design reviews aligning firmware, driver, runtime, and frameworks.
  • Engage with AMD customers and translate workload requirements into architectural direction.

Skills

Linux memory mgmt
GPU driver
Cache coherence
GPU networking
Scale-out networking
Infinity Fabric

Education

Bachelors/Masters in EE/CE/CS

Job description

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next‑generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.

Together, we advance your career.
THE ROLE

You will join the system software architecture team behind AMD Instinct™ accelerators, the GPUs powering some of the world's largest AI and HPC deployments. This role is focused on next‑generation, rack‑scale AI platforms in the MI400 class spanning the GPU, the node, and the scale‑up/scale‑out fabric that binds thousands of accelerators into a single training and inference system.

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next‑generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.

Together, we advance your career.
THE ROLE

You will join the system software architecture team behind AMD Instinct™ accelerators, the GPUs powering some of the world's largest AI and HPC deployments. This role is focused on next‑generation, rack‑scale AI platforms in the MI400 class spanning the GPU, the node, and the scale‑up/scale‑out fabric that binds thousands of accelerators into a single training and inference system.

THE PERSON

As a Systems Software Architect, you sit at the intersection of silicon, firmware, driver, runtime, and framework. You define how the software stack exposes and orchestrates the hardware so that AMD's largest customers can extract maximum performance, reliability, and utilization from their infrastructure.

Key Responsibilities
  • Own the end‑to‑end system software architecture for one or more MI400‑class subsystems for example GPU memory management, scheduling and queuing, RAS and serviceability, virtualization/partitioning (SR‑IOV), or the scale‑up/scale‑out interconnect software model.
  • Drive architecture across the stack: kernel‑mode driver (amdgpu/KFD), user‑mode runtime (ROCr/HSA), firmware interfaces, and the ROCm software platform, ensuring the layers compose cleanly and perform.
  • Partner with silicon and SoC architects during pre‑silicon definition to shape hardware/software interfaces, programming models, and register/firmware contracts before tape‑out.
  • Define the software strategy for multi‑GPU and rack‑scale topologies, including Infinity Fabric / UALink‑style interconnect, collective communication (RCCL), memory coherence, and address translation across the platform.
  • Establish architecture for reliability, availability, and serviceability at scale, error detection, containment, telemetry, recovery, and graceful degradation across large clusters.
  • Set direction on performance: identify bottlenecks in the launch path, memory subsystem, and communication path, and define the software mechanisms to close them.
  • Produce architecture specifications, reference designs, and design reviews that align firmware, driver, runtime, and framework teams onto a shared plan.
  • Act as a technical anchor across AMD and with strategic hyperscale and AI customers — translating their workload requirements into architectural direction and representing AMD in deep technical engagements.
Preferred Experience
  • Linux Memory Management and Heterogeneous Memory Management (HMM).
  • GPU / DRM driver development.
  • Cache coherence and memory consistency protocols.
  • GPU Networking technologies including including scale up transport, NVLink, UALink, RDMA, and peer‑direct.
  • Scale‑up and scale‑out networking, and communication collectives (e.g., RCCL/NCCL), MPI, or SHMEM.
  • Data‑center fabrics such as Infinity Fabric, UALink, Ultra Ethernet, InfiniBand, and RoCE, with topology‑aware software.
Desirable Experience
  • Direct experience with the ROCm stack, AMD Instinct, CUDA, or comparable GPU compute ecosystems.
  • Hands‑on experience developing or optimizing GPU compute kernels (HIP, CUDA, Triton, or assembly‑level tuning).
  • Experience with deep learning frameworks like PyTorch and TensorFlow in particular including framework integration, custom operators, and performance tuning on GPU backends.
  • Familiarity with the broader ML framework and compiler ecosystem (PyTorch, JAX, TensorFlow, ONNX, MLIR/compiler stacks).
  • Hands‑on experience with GPU compute technologies such as OpenCL and Vulkan.
  • Familiarity with AI/ML training and inference workloads (transformers, large‑scale distributed training, KV‑cache and memory pressure, inference serving).
  • Background in virtualization, multi‑tenancy, confidential computing, or cloud GPU provisioning.
  • Contributions to open‑source kernel, driver, or runtime projects.
Academic Credentials
  • Bachelor’s or Master’s in Electrical Engineer, Computer Engineering, Computer Science, or a closely related field
LOCATION:

Austin, TX

Santa Clara, CA

This role is not eligible for visa sponsorship.

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee‑based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Solutions Architect
Software Solutions Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 160,000
Software Solutions Architect
Software Solutions Architect

AMD • Austin (TX)

On-site
USD 130,000 - 160,000
Comprehensive benefits package
AI Platform Architect
AI Platform Architect

AMD • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Sr. Silicon Design Engineer: Hardware-Software Co-Design
Sr. Silicon Design Engineer: Hardware-Software Co-Design

AMD • Bellevue (WA)

Hybrid
USD 180,000 - 275,000
AMD benefits at a glance
AI Platform Architect
AI Platform Architect

Advanced Micro Devices • Santa Clara (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Platform Architect
AI Platform Architect

Advanced Micro Devices, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops
Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops

AMD • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Benefits at a glance
Lead Systems Management Architect
Lead Systems Management Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 260,000
Lead Systems Management Architect
Lead Systems Management Architect

Socket.dev • Austin (TX)

On-site
USD 170,000 - 250,000
Principal Enterprise AI/HPC GPU Systems Architect
Principal Enterprise AI/HPC GPU Systems Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 210,000 - 260,000