AI Instinct System Management Architect

AMD

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD in Santa Clara, CA is seeking an AI Instinct System Management Architect to define the architecture for system management across AI datacenter platforms. You will own reference designs and blueprints for deploying, integrating, and operating system management at rack and pod scale.

This role blends deep technical leadership with customer engagement, shaping security, observability, and manageability features that enable scalable AI workloads and reliable datacenter infrastructure.

Qualifications

  • Expert background in systems or platform software architecture for system management.
  • Deep expertise in BMC firmware stacks, telemetry, inventory, alerting, and management protocols.
  • Strong knowledge of DMTF standards (MCTP, PLDM, SPDM, Redfish), platform security, and management networking.
  • Experience with PCIe, CXL, NVMe interconnects and cluster schedulers (Kubernetes, Slurm).
  • Proven ability to combine technical leadership with customer engagement for scalable AI datacenter deployments.

Responsibilities

  • Define and drive architecture for system management across AI datacenter platforms.
  • Develop reference designs and blueprints for deploying OS management at scale.
  • Ensure compatibility with industry standards (Redfish, DMTF) and customer monitoring stacks.
  • Collaborate with customers to align with DCIM/ITSM environments and lead proofs of concepts.
  • Produce architectural collateral and telemetry baselines for customer readiness.

Skills

Systems architecture
BMC firmware
DMTF standards
PCIe/CXL/NVMe
Kubernetes/ Slurm
C/C++
Python/Go

Tools

Prometheus
Loki
ELK
OpenTelemetry

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.

Role Overview

The AI Instinct System Management Architect will define and drive the architecture for system management and observability across AMD’s AI datacenter platforms. This role spans firmware, operating systems, rack controllers, and orchestration layers to deliver highly manageable, composable, and secure infrastructure at rack and pod scale. It combines deep technical expertise with customer engagement, owning reference designs and blueprints for deploying, integrating, and operating system management at scale.

Key Priorities
  • Architect Scalable System Management: Develop a unified architecture for rack-scale and pod-scale management, ensuring seamless integration from device-level firmware to orchestration layers.
  • Industry Leadership: Benchmark against competitors, track emerging standards, and propose innovative features to position AMD as a leader in AI datacenter manageability.
  • Integration & Interoperability: Ensure compatibility with industry standards (Redfish, DMTF profiles) and customer monitoring stacks for observability and analytics.
  • Define and deliver manageability solutions (BMC/BSP, rack/pod controllers, APIs) ensuring coherent architecture from device to orchestration.
  • Develop standards-based interfaces and telemetry frameworks (Redfish, DMTF) for compute, storage, networking, and accelerators at scale.
  • Rack-Scale Lifecycle Management: Enable discovery, provisioning, firmware upgrades, and decommissioning workflows across racks and pods.
  • Collaborate with customers and partners to align with DCIM/ITSM environments, validate designs, and lead proof-of-concepts.
  • Produce architectural collateral and documentation, including reference designs, integration guides, and telemetry baselines for customer readiness.
  • Influence product strategy with customer-backed roadmaps optimized for AI workloads.
Required Qualifications
  • Expert background in systems or platform software architecture with focus on system management and server manageability.
  • Deep expertise in BMC firmware stacks, telemetry, inventory, alerting, and management protocols.
  • Strong knowledge of DMTF standards (MCTP, PLDM, SPDM, Redfish), platform security, and management networking.
  • Experience with PCIe, CXL, NVMe interconnects and cluster schedulers (Kubernetes, Slurm).
  • Proven ability to combine technical leadership with customer engagement for scalable AI datacenter deployments.
Preferred Qualifications
  • Experience with GPU/accelerator platforms
  • Familiarity with telemetry stacks (Prometheus, Loki, ELK, OpenTelemetry).
  • Knowledge of datacenter infrastructure components (racks, PDUs, power/thermal systems, fabric networking).
  • Contributions to open standards or open-source projects related to manageability or observability.
  • Strong programming skills in C/C++ and Python/Go; solid Linux systems experience.
Location:

Santa Clara, CA

This position is not eligible for Visa sponsorship

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Instinct System Management Architect
AI Instinct System Management Architect

Advanced Micro Devices • Santa Clara (CA), Northern (KY)

On-site
USD 180,000 - 240,000
AI Instinct System Management Architect
AI Instinct System Management Architect

Advanced Micro Devices, Inc. • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
Lead Systems Management Architect
Lead Systems Management Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 260,000
Lead Systems Management Architect
Lead Systems Management Architect

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 180,000 - 260,000
Lead Systems Management Architect
Lead Systems Management Architect

Socket.dev • Austin (TX)

On-site
USD 170,000 - 250,000
Director, AI Instinct Stack Architecture
Director, AI Instinct Stack Architecture

Socket.dev • Austin (TX)

Hybrid
USD 220,000 - 360,000
Benefits at a glance
Principal System Software Architect, AI/GPU Platforms
Principal System Software Architect, AI/GPU Platforms

AMD • Austin (TX)

On-site
USD 180,000 - 240,000
Lead Systems Management Architect
Lead Systems Management Architect

AMD • Austin (TX)

On-site
USD 180,000 - 240,000
AMD benefits
AI Platform Architect
AI Platform Architect

AMD • Santa Clara (CA)

On-site
USD 180,000 - 260,000
AI Platform Architect
AI Platform Architect

Advanced Micro Devices, Inc. • Santa Clara (CA)

On-site
USD 180,000 - 240,000