Lead Systems Management Architect

Socket.dev

Austin (TX)

On-site

USD 170,000 - 250,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD is seeking a Lead Systems Management Architect in Austin, TX, to lead OOB management for AMD Instinct accelerators and server platforms. You will shape the rack-scale management architecture and collaborate with BMC/firmware teams, ODM/OEMs, and customers to define robust interface standards.

You will drive architecture, conformance plans, and OpenBMC/Open standards involvement, mentoring engineers while advancing future roadmap discussions.

Qualifications

  • Experience delivering datacenter solutions for AI/HPC or cloud environments.
  • Ability to define architecture specifications and conformance plans.
  • Strong understanding of OOB management stack from DCIM to device protocols.

Responsibilities

  • Lead architecture and feature definition for OOB management of AMD Instinct accelerators and servers.
  • Define interface strategy using Redfish, PLDM, MCTP with AMD and OEM extensions.
  • Collaborate with BMC/firmware teams to specify Redfish schemas, inventory and eventing.
  • Manage firmware update flows, versioning, and recovery mechanisms for GPUs and servers.
  • Engage with customers/partners to translate requirements into concrete features and interfaces.

Skills

Platform management
Server manageability
BMC firmware
GPU/accelerator design
OOB management

Education

Bachelor’s or Master’s degree in CS/CE/EE

Tools

Redfish
PLDM
MCTP
OpenBMC

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.

THE ROLE:

AMD is seeking a Lead Systems Management Architect to lead key aspects of the architecture and feature definition for the management of AMD Instinct accelerators and associated server platforms. This is a highly visible technical leadership role contributing to the overall rack-scale management architecture. You will collaborate with fellow architects, internal engineering teams, AMD partners, and end customers to develop novel features and create solutions that support future AMD products and integration with customer datacenter infrastructure. A deep understanding of the entire out-of-band (OOB) ecosystem, from the DCIM SW layer down to the managed components, will be critical for success in this role.

THE PERSON:

The ideal candidate brings deep, hands‑on expertise across the full OOB management stack — comfortable reasoning about DCIM and orchestration software requirements at one end and getting into the details of BMC firmware, platform interfaces, and device‑level protocols at the other. You have a strong background in GPU/accelerator or server platform management and can translate that experience into clear architectural direction that works in practice, not just on paper. You are as effective writing a detailed interface specification or debugging a bring‑up issue as you are presenting a roadmap proposal to senior leadership or walking a customer through a platform integration. You work well in environments where requirements are still taking shape, build credibility through technical depth rather than title, and make the teams around you better.

KEY RESPONSIBILITIES:
  • Lead key aspects of architecture and feature definition for the OOB management of AMD Instinct GPU accelerators and associated server platforms, encompassing health monitoring, power and thermal management, firmware lifecycle, inventory, and error handling, while ensuring these capabilities integrate coherently into the broader rack-scale management architecture.
  • Define and contribute to the standards-based interface strategy for GPU and server platform manageability using DMTF Redfish, PLDM, MCTP, and related specifications, balancing standards compliance with AMD‑specific and OEM extension requirements.
  • Work with BMC and embedded firmware teams to define OOB management feature requirements, including Redfish schemas, sensor and inventory representations, eventing, firmware update flows, and debug workflows specific to GPU and server platform components.
  • Contribute to firmware management architecture for AMD Instinct accelerators and server platforms, covering in‑band and out‑of‑band update flows, versioning, dependency management, activation strategies, and recovery mechanisms.
  • Engage directly with end customers, AMD partners, and ODM/OEMs to understand datacenter integration requirements, DCIM and orchestration software expectations, and operational workflows, translating these into concrete feature and interface requirements and guiding partners through implementation.
  • Partner with GPU/SoC architects, board and system architects, firmware and software teams, security/RAS, and validation to translate architecture into production‑ready deliverables, and contribute to conformance and validation strategy for platform manageability.
  • Help shape future AMD Instinct platform roadmaps through customer engagement and field learnings, participate in relevant standards and open‑source communities including DMTF and OpenBMC, and mentor engineers and architects across the organization.
PREFERRED EXPERIENCE:
  • Expert‑level experience in platform management architecture, server manageability, BMC or embedded firmware, or GPU/accelerator platform design, including significant time in architect or technical leadership roles delivering solutions for datacenter, cloud, AI, or HPC environments.
  • Proven understanding of the full OOB management stack, from DCIM platforms and datacenter orchestration frameworks through BMC firmware down to device‑level management protocols, with the ability to reason clearly across every layer.
  • Deep knowledge of DMTF Redfish including schema design, OEM extension strategy, eventing, and update service; strong understanding of PLDM and MCTP for platform inventory, monitoring, control, and firmware update workflows; hands‑on experience with OpenBMC architecture and services is strongly preferred.
  • Experience with firmware security concepts including secure boot, root of trust, firmware signing, attestation, and SPDM, combined with a track record of producing architecture specifications, product requirements, conformance plans, and validation strategies that drive execution across internal teams and external partners.
  • Experience engaging directly with customers or ODM/OEM partners to gather requirements, present architecture proposals, and drive alignment on platform management capabilities; familiarity with AMD server or GPU platforms, AI/HPC system design, or OCP‑aligned rack architectures is a plus.
ACADEMIC CREDENTIALS:

Bachelor’s or Master’s degree (preferred) in Computer Science, Computer Engineering, Electrical Engineering, or equivalent

LOCATION:

Austin, TX

This role is not eligible for visa sponsorship.

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee‑based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Systems Management Architect
Lead Systems Management Architect

AMD • Austin (TX)

On-site
USD 180,000 - 240,000
AMD benefits
Lead Systems Management Architect
Lead Systems Management Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 260,000
Lead Systems Management Architect
Lead Systems Management Architect

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 180,000 - 260,000
Systems Application Engineer
Systems Application Engineer

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 130,000 - 180,000
Senior Manager, GPU Application Engineering
Senior Manager, GPU Application Engineering

AMD • Santa Clara (CA)

On-site
USD 130,000 - 160,000
Principal Platform System Engineering Lead
Principal Platform System Engineering Lead

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 160,000 - 260,000
AI Instinct System Management Architect
AI Instinct System Management Architect

Advanced Micro Devices • Santa Clara (CA), Northern (KY)

On-site
USD 180,000 - 240,000
Senior Manager, GPU Application Engineering
Senior Manager, GPU Application Engineering

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Software Solutions Architect
Software Solutions Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 160,000
Principal Enterprise AI/HPC GPU Systems Architect
Principal Enterprise AI/HPC GPU Systems Architect

AMD • Austin (TX)

On-site
USD 180,000 - 240,000
AMD benefits at a glance