AI Instinct System Management Architect

Advanced Micro Devices

Santa Clara, Northern (CA, KY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Advanced Micro Devices in Santa Clara, CA seeks an AI Instinct System Management Architect to define scalable system management across AI datacenter platforms. You will own reference designs spanning firmware to orchestration layers, delivering secure, manageable infrastructure at rack and pod scale.

The role requires deep expertise in BMC firmware, DMTF standards, PCIe/CXL/NVMe, and experience with Kubernetes or Slurm. This position is based in Santa Clara, CA and is not visa-eligible.

Qualifications

  • Expert background in systems or platform software architecture with focus on system management and server manageability.
  • Deep expertise in BMC firmware stacks, telemetry, inventory, alerting, and management protocols.
  • Strong knowledge of DMTF standards (MCTP, PLDM, SPDM, Redfish), platform security, and management networking.
  • Experience with PCIe, CXL, NVMe interconnects and cluster schedulers (Kubernetes, Slurm).
  • Proven ability to combine technical leadership with customer engagement for scalable AI datacenter deployments.

Responsibilities

  • Define and drive scalable system management architecture across rack-scale and pod-scale environments.
  • Lead integration of manageability features from device firmware to orchestration layers across datacenters.
  • Ensure compatibility with Redfish, DMTF profiles and customer monitoring stacks for observability.
  • Produce architectural collateral, reference designs, integration guides, and telemetry baselines.
  • Collaborate with customers and partners to align with DCIM/ITSM environments and validate designs.
  • Influence product strategy with customer-backed roadmaps for AI workloads.

Skills

System architecture
BMC firmware
DMTF standards
PCIe/CXL/NVMe
Kubernetes
Slurm
Customer engagement
Security
C/C++
Python/Go

Tools

Prometheus
Loki
ELK
OpenTelemetry

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.

Role Overview

The AI Instinct System Management Architect will define anddrivethe architecture for system management and observability across AMD’s AI datacenter platforms. This role spans firmware, operating systems, rack controllers, and orchestration layers to deliver highly manageable, composable, and secure infrastructure at rack and pod scale. It combines deep technical expertise with customer engagement, owning reference designs and blueprints for deploying, integrating, and operating system management at scale.

Key Priorities
  • Architect Scalable System Management: Develop a unified architecture for rack-scale and pod-scale management, ensuring seamless integration from device-level firmware to orchestration layers.
  • Industry Leadership: Benchmark against competitors, track emerging standards, and propose innovative features to position AMD as a leader in AI datacenter manageability.
  • Integration & Interoperability: Ensure compatibility with industry standards (Redfish, DMTF profiles) and customer monitoring stacks for observability and analytics.
  • Define and deliver manageability solutions (BMC/BSP, rack/pod controllers, APIs) ensuring coherent architecture from device to orchestration.
  • Develop standards-based interfaces and telemetry frameworks (Redfish, DMTF) for compute, storage, networking, and accelerators at scale.
  • Rack-Scale Lifecycle Management: Enable discovery, provisioning, firmware upgrades, and decommissioning workflows across racks and pods.
  • Collaborate with customers and partners to align with DCIM/ITSM environments, validate designs, and lead proof-of-concepts.
  • Produce architectural collateral and documentation, including reference designs, integration guides, and telemetry baselines for customer readiness.
  • Influence product strategy with customer-backed roadmaps optimized for AI workloads.
Required Qualifications
  • Expert background in systems or platform software architecture with focus on system management and server manageability.
  • Deep expertise in BMC firmware stacks, telemetry, inventory, alerting, and management protocols.
  • Strong knowledge of DMTF standards (MCTP, PLDM, SPDM, Redfish), platform security, and management networking.
  • Experience with PCIe, CXL, NVMe interconnects and cluster schedulers (Kubernetes, Slurm).
  • Proven ability to combine technical leadership with customer engagement for scalable AI datacenter deployments.
Preferred Qualifications
  • Experience with GPU/accelerator platforms
  • Familiarity with telemetrystacks (Prometheus, Loki, ELK, OpenTelemetry).
  • Knowledge of datacenter infrastructure components (racks, PDUs, power/thermal systems,fabric networking).
  • Contributions to open standards or open-source projects related to manageability or observability.
  • Strong programming skills in C/C++ and Python/Go; solid Linux systems experience.

Location: Santa Clara, CA

This position is not eligible for Visa sponsorship

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Instinct System Management Architect
AI Instinct System Management Architect

Advanced Micro Devices, Inc. • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
AI Instinct System Management Architect
AI Instinct System Management Architect

AMD • Santa Clara (CA)

On-site
USD 180,000 - 240,000
AI Platform Architect
AI Platform Architect

AMD • Austin (TX)

On-site
USD 180,000 - 260,000
Principal Platform System Engineering Lead
Principal Platform System Engineering Lead

AMD • Austin (TX)

On-site
USD 180,000 - 240,000
Principal Solutions Engineering – AI server/rack Infrastructure
Principal Solutions Engineering – AI server/rack Infrastructure

AMD • Seattle (WA)

On-site
USD 180,000 - 260,000
Director, AI Instinct Stack Architecture
Director, AI Instinct Stack Architecture

Advanced Micro Devices • Austin (TX)

On-site
USD 260,000 - 380,000
Principal Solutions Engineering – AI server/rack Infrastructure
Principal Solutions Engineering – AI server/rack Infrastructure

Advanced Micro Devices • Seattle (WA)

On-site
USD 180,000 - 240,000
Principal Platform System Engineering Lead
Principal Platform System Engineering Lead

Advanced Micro Devices • Austin (TX)

On-site
USD 150,000 - 230,000
AI Systems Security Architect
AI Systems Security Architect

Advanced Micro Devices • San Diego (CA)

On-site
USD 180,000 - 280,000
AMD benefits at a glance
Director, AI Instinct Stack Architecture
Director, AI Instinct Stack Architecture

AMD • Austin (TX)

On-site
USD 230,000 - 330,000