Technical Support Engineer

Advanced Micro Devices

Austin, Northern (TX, KY)

Hybrid

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

AMD benefits at a glance

Job summary

Advanced Micro Devices seeks a Data Center Operations Lead to oversee daily data center operations, guide a team of technicians, and provide advanced support for pre-production silicon and custom server platforms. This role blends people leadership with hands-on debugging and DCIM ownership to sustain secure, scalable operations.

You will collaborate with engineering stakeholders, drive improvements, and manage documentation, procedures, and runbooks.

Qualifications

  • 7+ years of experience in data center operations, server hardware support, hardware validation, or related technical environments.
  • Prior experience leading or supervising technical teams.
  • Strong knowledge of x86 server architecture, including CPUs, memory, PCIe, storage, networking, power, and thermal subsystems.
  • Hands-on experience with hardware and platform debugging tools, including BMC/IPMI/Redfish, serial consoles, and POST diagnostics.
  • Experience with DCIM or related asset, capacity, and infrastructure management systems.
  • Linux command-line proficiency and foundational networking knowledge.
  • Ability to work onsite in Austin, TX and support occasional after-hours maintenance and escalation activities.

Responsibilities

  • Lead, mentor, and develop a team of data center technicians through coaching, training, and performance management.
  • Manage staffing, schedules, PTO coverage, and escalation support to maintain operational service levels.
  • Partner with hiring managers on recruiting, interviewing, onboarding, and training new employees and contractors.
  • Establish and enforce standards for cabling, labeling, rack builds, ESD handling, safety, and lab cleanliness.
  • Translate engineering priorities into actionable work plans and serve as the primary escalation point for technical and operational issues.
  • Provide senior-level support for bring-up, validation, and debug of pre-production AMD silicon and custom server platforms.
  • Lead root-cause analysis of platform failures, including boot, POST, memory, PCIe, thermal, power, and firmware-related issues.
  • Utilize board-level and platform debugging techniques, including serial consoles, BMC/IPMI/Redfish, POST diagnostics, and lab instrumentation.
  • Perform BIOS, BMC, CPLD, and firmware updates, recovery, and image restoration.
  • Partner with silicon, firmware, validation, and platform engineering teams to drive issues through resolution.
  • Maintain secure handling, tracking, and disposition of pre-release hardware and sensitive assets.
  • Create and maintain runbooks, knowledge articles, and technical documentation.
  • Maintain ownership of DCIM data, including asset tracking, rack elevations, connectivity mapping, and capacity management.
  • Conduct audits and reconciliation activities to ensure asset and infrastructure accuracy.
  • Monitor data center power, cooling, environmental conditions, and redundancy health, proactively addressing risks.
  • Support capacity planning for space, power, cooling, and network resources.
  • Coordinate maintenance activities, vendor support, facility projects, and change-control processes.
  • Maintain operational documentation, diagrams, procedures, and disaster recovery plans.
  • Lead rack-and-stack, cabling, hardware deployment, relocation, and decommissioning projects.
  • Manage inventory, spare parts, asset tracking, and procurement support.
  • Drive process improvements, automation, and operational efficiencies.
  • Ensure compliance with AMD safety, security, environmental, and export control requirements.
  • Track and report operational metrics, team performance, incidents, and capacity trends.

Job description

The Position

AMD is seeking a Data Center Operations Lead to oversee daily data center operations, lead a team of technicians, and provide advanced support for pre-production silicon and custom server platforms. This role combines people leadership, hands-on platform debugging, DCIM ownership, and operational excellence to ensure a secure, efficient, and scalable environment supporting AMD engineering and validation teams.

The Person

You are a hands-on technical leader with experience managing teams in data center, server hardware, or lab environments. You excel at troubleshooting complex platform issues, driving operational improvements, and partnering with engineering stakeholders to deliver results. You bring strong organizational skills, a customer-focused mindset, and the ability to balance strategic planning with day-to-day execution.

Responsibilities
Team Leadership & Operations
  • Lead, mentor, and develop a team of data center technicians through coaching, training, and performance management.
  • Manage staffing, schedules, PTO coverage, and escalation support to maintain operational service levels.
  • Partner with hiring managers on recruiting, interviewing, onboarding, and training new employees and contractors.
  • Establish and enforce standards for cabling, labeling, rack builds, ESD handling, safety, and lab cleanliness.
  • Translate engineering priorities into actionable work plans and serve as the primary escalation point for technical and operational issues.
Advanced Platform Bring-Up & Debug
  • Provide senior-level support for bring-up, validation, and debug of pre-production AMD silicon and custom server platforms.
  • Lead root-cause analysis of platform failures, including boot, POST, memory, PCIe, thermal, power, and firmware-related issues.
  • Utilize board-level and platform debugging techniques, including serial consoles, BMC/IPMI/Redfish, POST diagnostics, and lab instrumentation.
  • Perform BIOS, BMC, CPLD, and firmware updates, recovery, and image restoration.
  • Partner with silicon, firmware, validation, and platform engineering teams to drive issues through resolution.
  • Maintain secure handling, tracking, and disposition of pre-release hardware and sensitive assets.
  • Create and maintain runbooks, knowledge articles, and technical documentation.
DCIM & Facilities Management
  • Maintain ownership of DCIM data, including asset tracking, rack elevations, connectivity mapping, and capacity management.
  • Conduct audits and reconciliation activities to ensure asset and infrastructure accuracy.
  • Monitor data center power, cooling, environmental conditions, and redundancy health, proactively addressing risks.
  • Support capacity planning for space, power, cooling, and network resources.
  • Coordinate maintenance activities, vendor support, facility projects, and change-control processes.
  • Maintain operational documentation, diagrams, procedures, and disaster recovery plans.
Continuous Improvement
  • Lead rack-and-stack, cabling, hardware deployment, relocation, and decommissioning projects.
  • Manage inventory, spare parts, asset tracking, and procurement support.
  • Drive process improvements, automation, and operational efficiencies.
  • Ensure compliance with AMD safety, security, environmental, and export control requirements.
  • Track and report operational metrics, team performance, incidents, and capacity trends.
Qualifications
  • 7+ years of experience in data center operations, server hardware support, hardware validation, or related technical environments.
  • Prior experience leading or supervising technical teams.
  • Strong knowledge of x86 server architecture, including CPUs, memory, PCIe, storage, networking, power, and thermal subsystems.
  • Hands-on experience with hardware and platform debugging tools, including BMC/IPMI/Redfish, serial consoles, and POST diagnostics.
  • Experience performing BIOS, BMC, and firmware updates and recovery procedures.
  • Experience with DCIM or related asset, capacity, and infrastructure management systems.
  • Understanding of data center power distribution, redundancy, cooling, and environmental monitoring.
  • Linux command-line proficiency and foundational networking knowledge.
  • Strong written and verbal communication skills with demonstrated technical documentation experience.
  • Ability to lift up to 50 lbs., work on ladders or lifts, and perform physical duties in a data center environment.
  • Ability to work onsite in Plano, TX and support occasional after-hours maintenance and escalation activities.
Preferred
  • Experience supporting pre-production or engineering-sample silicon programs.
  • Familiarity with AMD EPYC™, AMD Instinct™, or comparable server and accelerator platforms.
  • Experience with Sunbird dcTrack or similar enterprise DCIM solutions.
  • Experience with ServiceNow, Jira, ChangeGear, or other ITSM/change management platforms.
  • Scripting experience with Python, Bash, or PowerShell for automation and operational efficiency.
  • Experience supporting colocation facilities or multi-site data center environments.
Education
  • Bachelor's degree in Information Technology, Computer Engineering, Electrical Engineering, Computer Science, or a related field preferred.
  • Equivalent combination of education, military service, technical certifications, and relevant industry experience will also be considered.
Location:

Austin, Tx

This role is not eligible for visa sponsorship.

#LI-DNI

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s "Responsible AI Policy" is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Field Applications Engineer, Server Datacenter - Dell
Field Applications Engineer, Server Datacenter - Dell

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 160,000
Field Applications Engineer, Server Datacenter - Dell
Field Applications Engineer, Server Datacenter - Dell

CareerArc • Austin (TX)

On-site
USD 140,000 - 190,000
Field Applications Engineer, Server Datacenter - Dell
Field Applications Engineer, Server Datacenter - Dell

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
Data Center Technician
Data Center Technician

AMD • Boydton (VA)

On-site
USD 45,000 - 65,000
AMD benefits at a glance
Field Applications Engineer, Server Datacenter - HPE
Field Applications Engineer, Server Datacenter - HPE

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 150,000
Benefits at a glance
Field Applications Engineer, Server Datacenter - HPE
Field Applications Engineer, Server Datacenter - HPE

CareerArc • Austin (TX)

On-site
USD 110,000 - 160,000
North America Sr Solutions Architect/FAE - Channel - West U.S
North America Sr Solutions Architect/FAE - Channel - West U.S

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 160,000
North America Sr Solutions Architect/FAE - Channel - West U.S
North America Sr Solutions Architect/FAE - Channel - West U.S

AMD • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Technologist
Senior Technologist

AMD • Secaucus (NJ)

On-site
USD 65,000 - 95,000
AI /HPC Data Center Lab Engineer
AI /HPC Data Center Lab Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 90,000 - 120,000