DCGPU Platform System Manager

AMD

Austin (TX)

On-site

USD 130,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

AMD in Austin, TX is seeking a Data Center Platform Engineering Group Manager to lead a high-availability data center onsite with 1000+ systems. You will manage engineers and technicians, drive deployment/availability, and report status in executive settings.

Ideal candidates have prior GPU data center experience, strong capacity planning, and hands-on problem-solving skills to keep the fleet running smoothly.

Qualifications

  • Bachelor's degree in engineering or equivalent.
  • Experience managing a GPU data center with 500+ systems.
  • Strong on-site leadership and daily operations oversight.
  • Proven ability to communicate status to executives and stakeholders.

Responsibilities

  • On-site management of a large data center with diverse platforms.
  • Direct supervision of data center engineers and technicians.
  • Lead daily standups focused on system availability and issue resolution.
  • Root-cause analysis and debugging methodology mapping.
  • Provide leadership input to drive improvements and initiatives.

Skills

GPU data center
Data center management
Executive communication
Leadership
Capacity planning
Root cause analysis

Education

Bachelor's degree in engineering

Job description

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.

Together, we advance your career.
The Role

The Data Center Platform Engineering Group (DPEG) Manager that will be the primary on-site leader for a data center in north Austin, Texas. This individual will need to be on-site daily driving day-to-day activities that will include the deployment and availability of highly complex AMD Instinct platforms that will be the backbone of the work required to release state-of-the-art technologies. This manager will have personnel responsibilities and be required to manage the work required to run a data center with well over 1000 systems.

The Person

Experienced, self-motivated individual that has previously managed a data center preferably in the GPU system space w/500+ systems/platform. Person should be able to communicate updates on the state of the fleet, ensure the team is working on deployment/availability of the fleet, as well be able to solve technical issues that arise.

Key Responsibilities
  • On-site management of a data center site with a vast array of differing platforms/system
  • Personnel and work assigned management of data center engineers and technicians
  • Clear / Concise communication in open daily meetings on “high attention” given to system availability
  • Knowledge in requirements for root causing issues and understanding/mapping debug methodologies
  • Provide leadership input/recommendations for improvements / help drive organizational initiatives
Preferred Experience
  • Experience managing GPU data center employees, systems, day-to-day activities
  • Knowledge in solving issues around capacity planning, power, thermal, networking, clustering
  • Executive-level focused communication
  • Management experience in GPU data centers that have a vast array of different systems/platforms
  • Co-work with external stakeholders/vendors and ability to openly communicate/drive issues
  • Collaborate with internal stakeholders on root causing issues, driving issue meetings
  • Hands-on experience with day-to-day issues that arise in a data center (capacity, network, power, thermal)
  • Leadership and communication skills that require presentations in executive forum
  • Ability to clearly articulate the work by the team (ins/outs), needs, and recruit necessary skills
Academic Credentials
  • Bachelors degree in engineering
This role is not eligible for visa sponsorship.

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

DCGPU Platform System Manager
DCGPU Platform System Manager

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
Data Center Engineer
Data Center Engineer

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
Senior Datacenter Platform/Debug Engineer
Senior Datacenter Platform/Debug Engineer

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
AI /HPC Data Center Lab Engineer
AI /HPC Data Center Lab Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 90,000 - 120,000
Data Center Engineer
Data Center Engineer

Advanced Micro Devices • Town of Texas (WI)

On-site
USD 120,000 - 180,000
AMD benefits
Senior Datacenter Platform/Debug Engineer
Senior Datacenter Platform/Debug Engineer

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 110,000 - 160,000
AMD benefits at a glance
Senior Manager, GPU Application Engineering
Senior Manager, GPU Application Engineering

AMD • Santa Clara (CA)

On-site
USD 130,000 - 160,000
Director, Product Management - DC GPU
Director, Product Management - DC GPU

AMD • Santa Clara (CA)

On-site
USD 180,000 - 220,000
Comprehensive benefits
Flexible working hours
Career development opportunities
Technical Program Manager - AI/HPC Data Center Lab
Technical Program Manager - AI/HPC Data Center Lab

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 140,000 - 200,000
AMD benefits at a glance
Senior Manager, GPU Application Engineering
Senior Manager, GPU Application Engineering

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 240,000