Senior Systems Engineer

Fujitsu Limited

Singapore

On-site

SGD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Fujitsu Limited is seeking a seasoned HPC administrator to manage the day-to-day operations of a high-performance computing environment in Singapore. You will ensure high availability, stability, and peak performance while administering HPCM for cluster provisioning, monitoring, and lifecycle management.

You will also optimize HPC storage, deploy and maintain the PBS Professional workload manager, and troubleshoot the Slingshot interconnect to maximize efficiency.

Qualifications

  • 5–7 years of hands-on HPC administration experience in enterprise, research, or academic settings.
  • Familiarity with configuration management and automation tools is advantageous.
  • Understanding of GPU computing technologies (CUDA, ROCm) and accelerator-based HPC environments is a plus.

Responsibilities

  • Manage day-to-day operations of the HPC environment ensuring high availability, stability, performance and reliability.
  • Administer and maintain HPCM for cluster provisioning, monitoring, health management, software deployment and lifecycle management.
  • Manage and optimize compute infrastructure delivering high performance and efficient resource utilization.
  • Administer and optimize the Lustre parallel file system with large storage capacity and high I/O performance.
  • Configure, administer and maintain the PBS Professional workload manager including queues and scheduling policies.
  • Manage and troubleshoot the Slingshot interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Collaborate with infrastructure, storage, networking and application teams to support HPC deployments, upgrades and production operations.
  • Provide technical support to researchers, scientists and engineers and assist users in optimizing HPC applications.

Skills

HPC administration
Ansible
xCAT
Bright Cluster Manager
Infrastructure-as-Code
CUDA/ROCm GPU computing

Tools

HPE Cluster Manager (HPCM)
PBS Professional
HPE Slingshot interconnect
ClusterStor Lustre
IBM Storage Scale (GPFS)

Job description

Location Flexibility: Primary Location Only

Posting Start Date: 9/21/26

Responsibilites:
  • Manage the day-to-day operations of the HPE Cray EX supercomputing environment, ensuring high availability, stability, performance, and reliability of HPC services.
  • Administer and maintain HPE Cluster Manager (HPCM) for cluster provisioning, monitoring, health management, software deployment, and lifecycle management.
  • Manage AMD-based HPE Cray EX compute infrastructure delivering up to 10 PFLOPS of computational performance, ensuring optimal resource utilization and system efficiency.
  • Administer and optimize HPE ClusterStor Lustre parallel file system with over 10 PB of storage capacity, ensuring high-performance I/O, data integrity, and storage availability.
  • Manage IBM Storage Scale (formerly GPFS) parallel file system with over 15 PB of storage capacity, including performance tuning, capacity planning, and filesystem maintenance.
  • Configure, administer, and maintain the PBS Professional workload manager, including queue configuration, scheduling policies, fair-share management, resource allocation, and job troubleshooting.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Provide technical support to a diverse community of researchers, scientists, engineers, and academic users by troubleshooting application, storage, scheduler, and system-related issues.
  • Assist users in optimizing HPC applications through performance analysis, job scheduling best practices, parallel computing techniques, and efficient resource utilization.
  • Conduct user onboarding sessions, technical workshops, and training programs on HPC environment usage, job submission, parallel file systems, and cluster best practices.
  • Perform software installation, upgrades, patch management, and validation for HPC operating systems, middleware, compilers, MPI libraries, and scientific applications.
  • Develop and maintain automation scripts using Shell, Python, or similar scripting languages to streamline system administration, monitoring, reporting, and operational tasks.
  • Maintain comprehensive operational documentation, standard operating procedures (SOPs), architecture diagrams, and technical knowledge base articles. Participate in incident response, planned maintenance activities, disaster recovery exercises, and root cause analysis to ensure continuous improvement of HPC infrastructure.
  • Ensure adherence to security policies, operational standards, and best practices while maintaining a secure and highly available HPC environment.
  • Continuously evaluate emerging HPC technologies and recommend improvements to enhance system performance, scalability, reliability, and operational efficiency.
Requirements:
  • 5–7 years of hands-on experience administering High Performance Computing (HPC) environments in enterprise, research, or academic organizations.
  • Familiarity with configuration management and automation tools such as Ansible, xCAT, Bright Cluster Manager, or Infrastructure-as-Code solutions is an advantage.
  • Understanding of GPU computing technologies (NVIDIA CUDA, AMD ROCm) and accelerator-based HPC environments is an added advantage.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Engineer
Senior Systems Engineer

fujitsu asia pte ltd • Singapore

On-site
SGD 90,000 - 130,000
Senior Systems Engineer
Senior Systems Engineer

Fujitsu • Singapore

On-site
SGD 120,000 - 180,000
System Engineer(HPC)
System Engineer(HPC)

RAPSYS TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 70,000 - 110,000
HPC System Administrator, System, NSCC
HPC System Administrator, System, NSCC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 60,000 - 90,000
HPC Storage Engineer (System), NSCC
HPC Storage Engineer (System), NSCC

A*STAR Research Entities (A*STAR) • Singapore

On-site
SGD 90,000 - 125,000
Systems Engineer
Systems Engineer

Fujitsu • Singapore

On-site
SGD 90,000 - 120,000
Server Engineer(AI Cluster/GPU)
Server Engineer(AI Cluster/GPU)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 180,000
Senior HPC Systems Engineer — Exascale Compute & Storage
Senior HPC Systems Engineer — Exascale Compute & Storage

Fujitsu • Singapore

On-site
SGD 120,000 - 180,000
Senior HPC Systems Engineer: Clusters & Storage
Senior HPC Systems Engineer: Clusters & Storage

Fujitsu Limited • Singapore

On-site
SGD 120,000 - 180,000
HPC Middleware Engineer, System, NSCC
HPC Middleware Engineer, System, NSCC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 70,000 - 90,000