Infrastructure Systems Engineer

AMD

Penang

Hybrid

MYR 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

AMD benefits at a glance

Job summary

AMD is seeking an Infrastructure Systems Engineer in Penang to provide advanced operational support for compute, virtualization, and storage services. The role acts as a NOC escalation point, drives durable improvements in monitoring and runbooks, and coordinates with platform and domain teams on incidents.

You will bridge day-to-day operations with infrastructure engineering, converting recurring issues into durable improvements and ensuring reliable platform delivery.

Qualifications

  • Bachelor's degree in Computer Science, IT, or engineering or equivalent experience.
  • Hands-on Linux, virtualization, storage, and automation fundamentals.
  • Experience with data-center infrastructure and incident response.

Responsibilities

  • Provide L2 infrastructure systems support for Linux hosts, bare-metal, virtualization, and storage dependencies.
  • Validate findings, reproduce issues, identify failure domains, and execute approved recovery actions.
  • Investigate incidents using logs, metrics, telemetry, and dependency maps.
  • Support server lifecycle activities including provisioning, firmware readiness, and decommissioning.
  • Troubleshoot Linux services, filesystems, permissions, and connectivity dependencies.
  • Support bare-metal provisioning and orchestration (PXE, Redfish, MAAS).
  • Support virtualization platforms like Xen/VMware and assess host health and recovery.
  • Coordinate with storage teams for platform-level issues (NetApp, Pure, Weka, Ceph).
  • Collaborate with cross-domain teams to isolate incidents and maintain ownership.
  • Create escalation packages with impact, status, and recommended steps.

Skills

Linux administration
Virtualization
Storage administration
Automation
Troubleshooting
Incident management
Networking basics
Scripting (Python/Shell)
Hardware health/BIOS/Redfish
PXE/MAAS/Ironic provisioning

Education

Bachelor's degree in CS/IT/Engineering
IT service-management certification beneficial

Tools

MAAS
PXE
Redfish
Ironic
TF/CI-CD tooling
NetApp/Pure/Weka storage

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.

THE ROLE:

We are seeking an Infrastructure Systems Engineer to provide advanced operational support for the compute, virtualization, storage, operating-system, and bare-metal provisioning services used by AMD Fleet Services. The role serves as an escalation point for the IOQ Modern NOC, validates first-line findings, restores service through authorized actions, and coordinates with platform and domain engineering teams when incidents require deeper expertise.

This engineer will bridge day-to-day NOC operations and infrastructure engineering by converting recurring issues into durable improvements to monitoring, automation, standard operating procedures, runbooks, and platform reliability.

THE PERSON:

You are a hands-on infrastructure engineer with strong Linux, compute, virtualization, storage, and automation fundamentals. You troubleshoot methodically across hardware and software boundaries, use evidence to isolate failure domains, and communicate clear ownership, risk, and escalation requirements during high-impact incidents.

KEY RESPONSIBILITIES:
  • Provide L2 infrastructure systems support for in-scope Fleet Services environments, including Linux hosts, bare-metal systems, virtualization platforms, storage dependencies, and supporting management services.
  • Validate L1 findings, reproduce symptoms where practical, identify the affected failure domain, and execute approved recovery actions within assigned access and change authority.
  • Investigate incidents using system logs, service state, hardware telemetry, operating-system metrics, configuration data, dependency maps, and infrastructure management tools.
  • Support server lifecycle activities such as provisioning, configuration validation, firmware and operating-system readiness, capacity checks, maintenance preparation, and decommissioning where assigned.
  • Troubleshoot Linux services, processes, filesystems, permissions, package or configuration issues, resource contention, and connectivity dependencies that affect infrastructure availability.
  • Support bare-metal provisioning and orchestration workflows, including PXE, Redfish, MAAS, or comparable technologies used by Fleet Services.
  • Support virtualization and compute platforms such as Xen, VMware, or comparable environments, including host health, guest dependencies, resource allocation, and recovery validation.
  • Assess storage-related symptoms and dependencies, gather evidence, and coordinate with storage engineering for platform-level issues involving systems such as NetApp, Pure, Weka, or comparable technologies.
  • Partner with network, cloud, identity, facilities, platform, and service owners to isolate cross-domain incidents and maintain clear technical ownership through resolution.
  • Create complete escalation packages with impact, timeline, affected assets, evidence collected, actions attempted, current status, risks, and recommended next steps.
  • Participate in major-incident response, technical bridges, recovery validation, postmortems, and corrective-action tracking for infrastructure-related events.
  • Develop and maintain SOPs, runbooks, troubleshooting guides, validation checks, and knowledge articles that enable safe and consistent L1 execution.
  • Automate repeatable diagnostics, health checks, evidence collection, configuration validation, and approved remediation using Python, shell, Ansible, APIs, or comparable tools.
  • Improve infrastructure observability by defining actionable metrics, alerts, dashboards, service-health indicators, and evidence requirements with IOQ and engineering partners.
  • Identify recurring incidents, monitoring gaps, configuration drift, capacity risks, and reliability improvements, and drive them to the appropriate owner with measurable follow-through.
  • Maintain accurate Jira records, change references, operational documentation, and follow-the-sun handoffs throughout the incident lifecycle.
PREFERRED EXPERIENCE:
  • Experience supporting large-scale Linux, data center, GPU/HPC, bare-metal, private-cloud, or hybrid production infrastructure.
  • Strong Linux administration and troubleshooting across services, filesystems, networking, permissions, performance, and logs.
  • Experience with server hardware, firmware, BIOS, out-of-band management, Redfish, and hardware-health telemetry.
  • Experience with bare-metal provisioning or lifecycle platforms such as MAAS, PXE, Ironic, Tinkerbell, or comparable tooling.
  • Working knowledge of enterprise storage, including NetApp, Pure, Weka, Ceph, Lustre, NFS, or block storage.
  • Automation experience with Python, shell, Ansible, Terraform, APIs, CI/CD, or configuration-management systems.
  • Ability to lead cross-team troubleshooting and distinguish immediate recovery from longer-term corrective action.
ACADEMIC CREDENTIALS:
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field preferred; equivalent relevant experience considered.
  • Linux, virtualization, cloud, storage, networking, automation, or IT service-management certification is beneficial.
LOCATION:

Penang, Malaysia

#LI-KL1

#LI-Hybrid

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Infrastructure Systems Engineer
Infrastructure Systems Engineer

Advanced Micro Devices, Inc. • Penang

Hybrid
MYR 120,000 - 180,000
AMD benefits at a glance
Infrastructure Systems Engineer
Infrastructure Systems Engineer

Advanced Micro Devices, Inc. • Bayan Lepas

On-site
MYR 90,000 - 150,000
Infrastructure Platform Engineer
Infrastructure Platform Engineer

Advanced Micro Devices, Inc. • Penang

Hybrid
MYR 120,000 - 180,000
Infrastructure Platform Engineer
Infrastructure Platform Engineer

AMD • Penang

Hybrid
MYR 180,000 - 240,000
AMD benefits
Infrastructure Platform Engineer
Infrastructure Platform Engineer

Advanced Micro Devices, Inc. • Bayan Lepas

Hybrid
MYR 180,000 - 240,000
AMD benefits at a glance
Graduate Trainee - IT Engineer
Graduate Trainee - IT Engineer

Advanced Micro Devices • Malaysia

On-site
MYR 42,000 - 68,000
AMD benefits at a glance
Staff Post Silicon Validation Engineer (Automation)
Staff Post Silicon Validation Engineer (Automation)

Advanced Micro Devices • Penang

On-site
MYR 180,000 - 240,000
Staff Post Silicon Validation Engineer (Automation)
Staff Post Silicon Validation Engineer (Automation)

AMD • George Town

On-site
MYR 120,000 - 180,000
Principal Datacenter Platform Validation Lab Engineer
Principal Datacenter Platform Validation Lab Engineer

Advanced Micro Devices • Penang

On-site
MYR 89,000 - 179,000
Senior Post Silicon Validation Engineer (Automation)
Senior Post Silicon Validation Engineer (Automation)

Advanced Micro Devices • Penang

On-site
MYR 180,000 - 240,000
AMD benefits