Senior Systems Engineer

Fujitsu

Singapore

On-site

SGD 120,000 - 180,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Fujitsu in Singapore is seeking an experienced HPC systems administrator to manage the day-to-day operations of a high-performance computing environment. You will handle cluster provisioning, health monitoring, software deployment, and lifecycle management, ensuring high availability and performance for researchers and engineers.

You will work with IBM/GPU accelerated infrastructure, Slingshot interconnects, Lustre file systems, and PBS workload scheduling, collaborating with cross-functional

Qualifications

  • 5–7 years hands-on HPC administration experience.
  • Experience with Ansible, xCAT, or Bright Cluster Manager is a plus.
  • Understanding GPU computing (CUDA/ROCm) is advantageous.

Responsibilities

  • Manage day-to-day operations of the HPC environment, ensuring high availability and performance.
  • Administer and maintain cluster management, software deployment, and lifecycle management.
  • Configure, monitor, and optimize parallel file systems and interconnects.
  • Provide technical support to researchers and engineers; assist in performance tuning.
  • Collaborate with cross-functional teams for deployments, upgrades, and maintenance.

Skills

HPC administration
Cluster management
Performance tuning
Shell scripting
Python

Tools

Ansible
xCAT
Bright Cluster Manager
Infrastructure-as-Code

Job description

Job Location: Singapore

Location Flexibility: Primary Location Only

Req Id: 12418

Posting Start Date: 9/21/26

Responsibilites
  • Manage the day-to-day operations of the HPE Cray EX supercomputing environment, ensuring high availability, stability, performance, and reliability of HPC services.
  • Administer and maintain HPE Cluster Manager (HPCM) for cluster provisioning, monitoring, health management, software deployment, and lifecycle management.
  • Manage AMD-based HPE Cray EX compute infrastructure delivering up to 10 PFLOPS of computational performance, ensuring optimal resource utilization and system efficiency.
  • Administer and optimize HPE ClusterStor Lustre parallel file system with over 10 PB of storage capacity, ensuring high-performance I/O, data integrity, and storage availability.
  • Manage IBM Storage Scale (formerly GPFS) parallel file system with over 15 PB of storage capacity, including performance tuning, capacity planning, and filesystem maintenance.
  • Configure, administer, and maintain the PBS Professional workload manager, including queue configuration, scheduling policies, fair-share management, resource allocation, and job troubleshooting.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Provide technical support to a diverse community of researchers, scientists, engineers, and academic users by troubleshooting application, storage, scheduler, and system-related issues.
  • Assist users in optimizing HPC applications through performance analysis, job scheduling best practices, parallel computing techniques, and efficient resource utilization.
  • Conduct user onboarding sessions, technical workshops, and training programs on HPC environment usage, job submission, parallel file systems, and cluster best practices.
  • Perform software installation, upgrades, patch management, and validation for HPC operating systems, middleware, compilers, MPI libraries, and scientific applications.
  • Develop and maintain automation scripts using Shell, Python, or similar scripting languages to streamline system administration, monitoring, reporting, and operational tasks.
  • Maintain comprehensive operational documentation, standard operating procedures (SOPs), architecture diagrams, and technical knowledge base articles. Participate in incident response, planned maintenance activities, disaster recovery exercises, and root cause analysis to ensure continuous improvement of HPC infrastructure.
  • Ensure adherence to security policies, operational standards, and best practices while maintaining a secure and highly available HPC environment.
  • Continuously evaluate emerging HPC technologies and recommend improvements to enhance system performance, scalability, reliability, and operational efficiency.
Requirements
  • 5–7 years of hands-on experience administering High Performance Computing (HPC) environments in enterprise, research, or academic organizations.
  • Familiarity with configuration management and automation tools such as Ansible, xCAT, Bright Cluster Manager, or Infrastructure-as-Code solutions is an advantage.
  • Understanding of GPU computing technologies (NVIDIA CUDA, AMD ROCm) and accelerator-based HPC environments is an added advantage.

Relocation Supported: No

Visa Sponsorship Approved: No

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Engineer
Senior Systems Engineer

Fujitsu Limited • Singapore

On-site
SGD 120,000 - 180,000
Senior Systems Engineer
Senior Systems Engineer

fujitsu asia pte ltd • Singapore

On-site
SGD 90,000 - 130,000
Systems Engineer
Systems Engineer

Fujitsu • Singapore

On-site
SGD 90,000 - 120,000
Senior HPC Systems Engineer — Exascale Compute & Storage
Senior HPC Systems Engineer — Exascale Compute & Storage

Fujitsu • Singapore

On-site
SGD 120,000 - 180,000
System Engineer(HPC)
System Engineer(HPC)

RAPSYS TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 70,000 - 110,000
Senior HPC Systems Engineer: Performance & Automation
Senior HPC Systems Engineer: Performance & Automation

fujitsu asia pte ltd • Singapore

On-site
SGD 90,000 - 130,000
Linux System Administrator (HPC)
Linux System Administrator (HPC)

OPUS IT Services Pte Ltd • Singapore

On-site
SGD 67,000 - 78,000
Senior HPC Systems Engineer: Clusters & Storage
Senior HPC Systems Engineer: Clusters & Storage

Fujitsu Limited • Singapore

On-site
SGD 120,000 - 180,000
HPC Performance Engineer,Frontier,NSCC
HPC Performance Engineer,Frontier,NSCC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 90,000 - 150,000
HPC Systems Engineer - Linux & Performance
HPC Systems Engineer - Linux & Performance

GOLDTECH RESOURCES PTE LTD • Singapore

On-site
SGD 80,000 - 100,000