Senior Systems Engineer

fujitsu asia pte ltd

Singapore

On-site

SGD 90,000 - 130,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Fujitsu Asia Pte Ltd is seeking an experienced HPC Administrator to manage and optimize our high-performance computing environment in Singapore. You will oversee cluster operations, storage systems, job scheduling, and interconnects while providing hands-on support to researchers and engineers.

You will work with HPCM, Lustre, GPFS, PBS, and Slingshot to ensure peak performance, reliability, and security. A strong scripting and automation mindset is essential for ongoing improvements.

Qualifications

  • 5-7 years hands-on experience administering HPC environments.
  • Familiarity with config/automation tools is a plus.
  • Understanding of GPU computing and accelerator HPC is a plus.

Responsibilities

  • Manage day-to-day operations of the HPC environment with high availability and reliability.
  • Administer HPE Cluster Manager for provisioning, monitoring, and software deployment.
  • Manage HPC compute infrastructure delivering high performance (up to 10 PFLOPS).
  • Administer and optimize Lustre and IBM Storage Scale parallel file systems.
  • Configure and maintain PBS Professional workload manager and job scheduling.
  • Troubleshoot Slingshot interconnect fabric and monitor cluster health.
  • Provide user support to researchers and engineers; train users on HPC usage.
  • Develop automation scripts to streamline administration and operations.
  • Maintain documentation, SOPs, and incident response playbooks.
  • Participate in maintenance, DR exercises, and security compliance.

Skills

HPC administration
System administration
Automation & scripting
Cluster management

Tools

Ansible
xCAT
Bright Cluster Manager
Infrastructure-as-Code

Job description

Responsibilities:
  • Manage the day-to-day operations of the HPE Cray EX supercomputing environment, ensuring high availability, stability, performance, and reliability of HPC services.
  • Administer and maintain HPE Cluster Manager (HPCM) for cluster provisioning, monitoring, health management, software deployment, and lifecycle management.
  • Manage AMD-based HPE Cray EX compute infrastructure delivering up to 10 PFLOPS of computational performance, ensuring optimal resource utilization and system efficiency.
  • Administer and optimize HPE ClusterStor Lustre parallel file system with over 10 PB of storage capacity, ensuring high-performance I/O, data integrity, and storage availability.
  • Manage IBM Storage Scale (formerly GPFS) parallel file system with over 15 PB of storage capacity, including performance tuning, capacity planning, and filesystem maintenance.
  • Configure, administer, and maintain the PBS Professional workload manager, including queue configuration, scheduling policies, fair-share management, resource allocation, and job troubleshooting.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems.
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability.
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations.
  • Provide technical support to a diverse community of researchers, scientists, engineers, and academic users by troubleshooting application, storage, scheduler, and system-related issues.
  • Assist users in optimizing HPC applications through performance analysis, job scheduling best practices, parallel computing techniques, and efficient resource utilization.
  • Conduct user onboarding sessions, technical workshops, and training programs on HPC environment usage, job submission, parallel file systems, and cluster best practices.
  • Perform software installation, upgrades, patch management, and validation for HPC operating systems, middleware, compilers, MPI libraries, and scientific applications.
  • Develop and maintain automation s using Shell, Python, or similar ing languages to streamline system administration, monitoring, reporting, and operational tasks.
  • Maintain comprehensive operational documentation, standard operating procedures (SOPs), architecture diagrams, and technical knowledge base articles. Participate in incident response, planned maintenance activities, disaster recovery exercises, and root cause analysis to ensure continuous improvement of HPC infrastructure.
  • Ensure adherence to security policies, operational standards, and best practices while maintaining a secure and highly available HPC environment.
  • Continuously evaluate emerging HPC technologies and recommend improvements to enhance system performance, scalability, reliability, and operational efficiency.
Requirements:
  • 5-7 years of hands-on experience administering High Performance Computing (HPC) environments in enterprise, research, or academic organizations.
  • Familiarity with configuration management and automation tools such as Ansible, xCAT, Bright Cluster Manager, or Infrastructure-as-Code solutions is an advantage.
  • Understanding of GPU computing technologies (NVIDIA CUDA, AMD ROCm) and accelerator-based HPC environments is an added advantage.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Systems Engineer
Senior Systems Engineer

Fujitsu Limited • Singapore

On-site
SGD 120,000 - 180,000
Senior Systems Engineer
Senior Systems Engineer

Fujitsu • Singapore

On-site
SGD 120,000 - 180,000
System Engineer(HPC)
System Engineer(HPC)

RAPSYS TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 70,000 - 110,000
HPC System Administrator, System, NSCC
HPC System Administrator, System, NSCC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 60,000 - 90,000
HPC Storage Engineer (System), NSCC
HPC Storage Engineer (System), NSCC

A*STAR Research Entities (A*STAR) • Singapore

On-site
SGD 90,000 - 125,000
Server Engineer(AI Cluster/GPU)
Server Engineer(AI Cluster/GPU)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 180,000
HPC System Admin
HPC System Admin

OPUS IT Services Pte Ltd • Singapore

On-site
SGD 60,000 - 90,000
Senior HPC Systems Engineer — Exascale Compute & Storage
Senior HPC Systems Engineer — Exascale Compute & Storage

Fujitsu • Singapore

On-site
SGD 120,000 - 180,000
HPC Middleware Engineer, System, NSCC
HPC Middleware Engineer, System, NSCC

A*STAR - Agency for Science, Technology and Research • Singapore

On-site
SGD 70,000 - 90,000
Senior HPC Systems Engineer: Clusters & Storage
Senior HPC Systems Engineer: Clusters & Storage

Fujitsu Limited • Singapore

On-site
SGD 120,000 - 180,000