Senior Datacenter Platform/Debug Engineer

Advanced Micro Devices, Inc.

Austin (TX)

On-site

USD 110,000 - 160,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

AMD benefits at a glance

Job summary

Advanced Micro Devices, Inc. (AMD) is seeking a hands-on Platform Systems Engineer to support deployment, availability, and operational success of AI and HPC infrastructure.

You will work across hardware, firmware, software, and datacenter teams to troubleshoot complex issues in large-scale compute environments. Ideal candidates bring strong troubleshooting and root cause analysis, excellent collaboration and documentation skills, and the ability to mentor junior engineers in a fast-paced

Qualifications

  • Hands-on systems engineer with deep technical troubleshooting across hardware, firmware and software layers.

Responsibilities

  • Support datacenter deployments and maintain availability and uptime of large-scale compute systems.
  • Perform system-level debugging and triage across hardware, firmware, software, and operating system layers.
  • Investigate and resolve complex platform issues impacting GPU and server infrastructure.
  • Support system bring-up, initialization, validation, and operational readiness activities.
  • Utilize industry-standard debug tools and diagnostic methods to identify root causes.
  • Provide technical leadership and guidance to junior engineers during troubleshooting.
  • Document debug methodologies, troubleshooting procedures, and best practices.
  • Collaborate with cross-functional engineering teams to drive issue resolution and continuous improvement.

Skills

Troubleshooting
Root cause analysis
Collaboration
Mentoring
Documentation
Linux
Python
Bash
Hardware debugging
System bring-up

Education

Bachelor's or Master's degree preferred in Computer Engineering, Electrical Engineering, Computer Science, or related

Tools

Debug tools

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou'redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger- technologythat moves the world forward.Join us and, together, we'll advance your career.

THE ROLE:

Join AMD's Datacenter Platform Engineering Group (DPEG) and help support the deployment, availability, and operational success of next-generation AI and HPC infrastructure. As a Platform Systems Engineer, you will work on cutting-edge GPU and server platforms, partnering with hardware, firmware, software, validation, and datacenter engineering teams to troubleshoot complex system issues and ensure reliable operation of large-scale compute environments.

This role offers the opportunity to work directly with advanced datacenter technologies, participate in system bring-up and deployment activities, and become a key contributor in resolving critical platform-level issues. Ideal candidates enjoy solving challenging technical problems, collaborating across multiple engineering disciplines, and making a direct impact on the success of AMD's datacenter infrastructure.

THE PERSON:

The ideal candidate is a hands-on systems engineer who enjoys deep technical troubleshooting and thrives in fast-paced datacenter environments. They possess strong analytical skills, can quickly isolate and resolve complex issues, and are comfortable working across hardware, firmware, and software layers of a system.

Successful candidates will demonstrate:

  • Strong troubleshooting and root cause analysis skills
  • Excellent communication and collaboration abilities
  • A proactive and self-driven approach to problem solving
  • Ability to mentor and guide junior engineers
  • Strong documentation and organizational skills
  • Comfort operating in highly technical and mission-critical environments
  • A passion for learning new technologies and solving complex engineering challenges
KEY RESPONSIBILITIES:
  • Support datacenter deployments and help maintain the availability and uptime of large-scale compute systems.
  • Perform system-level debugging and triage across hardware, firmware, software, and operating system layers.
  • Investigate and resolve complex platform issues impacting GPU and server infrastructure.
  • Support system bring-up, initialization, validation, and operational readiness activities.
  • Utilize industry-standard debug tools and diagnostic methods to identify root causes.
  • Provide technical leadership and guidance to junior engineers during troubleshooting activities.
  • Document debug methodologies, troubleshooting procedures, and best practices.
  • Collaborate with cross-functional engineering teams to drive issue resolution and continuous improvement.
PREFERRED EXPERIENCE:
  • System-level hardware, firmware, and software debugging experience
  • Datacenter, server, HPC, or AI infrastructure environments
  • Root cause analysis and triage of complex platform issues
  • GPU, PCIe, memory, retimer, networking, and system architecture knowledge
  • RAS (Reliability, Availability, Serviceability) concepts and methodologies
  • Linux operating system experience
  • Python, Bash, or similar scripting experience
  • Hands-on experience with industry-standard debug tools and diagnostics
  • Server bring-up, system initialization, and validation activities
  • Technical leadership, mentoring, and cross-functional collaboration
ACADEMIC CREDENTIALS:
  • Bachelor's or Master's degree preferred in Computer Engineering, Electrical Engineering, Computer Science, or a related technical discipline
LOCATION:

Rockdale, Texas | 100% Onsite

THIS ROLE IS NOT ELIGIBLE FOR VISA SUPPORT
LI-CS1

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lab Systems Engineer
Senior Lab Systems Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 160,000
AI /HPC Data Center Lab Engineer
AI /HPC Data Center Lab Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 90,000 - 120,000
Data Center Engineer
Data Center Engineer

AMD • United States

On-site
USD 120,000 - 170,000
Senior Lab Systems Engineer
Senior Lab Systems Engineer

Socket.dev • Austin (TX)

On-site
USD 120,000 - 180,000
Platform / System Debug Validation Engineer
Platform / System Debug Validation Engineer

AMD • Austin (TX)

On-site
USD 120,000 - 170,000
Platform / System Debug Validation Engineer
Platform / System Debug Validation Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
Data Center Engineer
Data Center Engineer

Advanced Micro Devices • Temple (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
AMD Benefits
Data Center Engineer
Data Center Engineer

AMD • Temple (TX)

On-site
USD 95,000 - 140,000
Data Center Engineer
Data Center Engineer

Socket.dev • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Product Development Engineer, Datacenter Systems (AI/HPC)
Product Development Engineer, Datacenter Systems (AI/HPC)

AMD • Austin (TX)

On-site
USD 110,000 - 170,000