Senior HPC Systems Engineer

The Massachusetts Green High Performance Computing Center Inc

Holyoke, Northern (MA, KY)

Hybrid

USD 120,000 - 163,000

Full time

17 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

The Massachusetts Green High Performance Computing Center Inc. in Holyoke, MA is seeking an HPC Systems Engineer to lead deployment, maintenance, and optimization of HPC clusters, storage systems, and networking for AI/ML workloads.

This hands-on role collaborates with researchers and engineers to ensure secure, efficient, and reliable infrastructure. Responsibilities include installing, configuring hardware/software, monitoring performance, backing up data, and developing documentation.

Qualifications

  • Bachelor's degree in CS/Engineering or related field.
  • Minimum 7 years of relevant experience.
  • Proven HPC System Engineer experience in HPC environments.
  • Deep knowledge of HPC architectures, Linux/Unix, cluster management.
  • Experience with MPI, job schedulers, and batch systems.
  • Storage expertise (Lustre/GPFS) and network file systems.
  • Containerization/virtualization (Docker, Kubernetes).
  • Strong troubleshooting skills.
  • Excellent communication with technical and non-technical stakeholders.

Responsibilities

  • Manage infrastructure, support, maintenance, and operations of systems, hardware, and software.
  • Install, configure, and maintain HPC hardware and software, including clusters, storage systems, and job schedulers.
  • Monitor system performance, identify bottlenecks, and implement optimizations to improve efficiency and resource utilization.
  • Provide technical support to end-users, including troubleshooting hardware and software issues.
  • Perform system backups and disaster recovery procedures to ensure data integrity and availability.
  • Develop and maintain system documentation, including installation guides, configuration files, and standard operating procedures.
  • Assist in the design and deployment of scalable HPC environments to meet the needs of scientific and engineering applications.
  • Collaborate with teams to identify and implement software and hardware upgrades.
  • Ensure security of HPC systems by applying patches, configuring firewalls, and performing regular audits.
  • Conduct performance benchmarking, diagnostics, and capacity planning to keep systems up to date with evolving needs.
  • Perform other duties as required.

Skills

HPC architectures
Linux/Unix
Cluster management
MPI / parallel programming
Job schedulers
Storage systems
Docker / Kubernetes
Python / Bash / Perl
Networking tools
Troubleshooting
Communication skills
Cloud HPC (AWS/GCP)

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

Nagios
Prometheus

Job description

Support the computing infrastructure behind the AI Computing Resource (AICR) that supports the Massachusetts AI Hub. This hands-on role will be responsible for deploying, maintaining, and optimizing HPC clusters, storage systems, and networking for AI/ML workloads. Join a collaborative, fast-paced team delivering critical infrastructure to some of the nation’s leading AI researchers.

Pay Range
$120,440 - $163,200
Position Overview
As a technical leader, the HPC Systems Engineer will manage and optimize high-performance computing environments, ensuring the smooth operation of complex clusters, and providing excellent support to end-users. Responsible for the installation, configuration, maintenance, and troubleshooting of HPC systems and infrastructure. Collaborate with researchers, engineers, and other IT professionals to ensure the HPC environment is secure, efficient, and reliable.
Principal Responsibilities
  • Manage infrastructure, support, maintenance, and operations of systems, hardware, and software.
  • Install, configure, and maintain HPC hardware and software, including clusters, storage systems, and job schedulers.
  • Monitor system performance, identify bottlenecks, and implement optimizations to improve efficiency and resource utilization.
  • Provide technical support to end-users, including troubleshooting hardware and software issues.
  • Perform system backups and disaster recovery procedures to ensure data integrity and availability.
  • Develop and maintain system documentation, including installation guides, configuration files, and standard operating procedures.
  • Assist in the design and deployment of scalable HPC environments to meet the needs of scientific and engineering applications.
  • Collaborate with teams to identify and implement software and hardware upgrades.
  • Ensure security of HPC systems by applying patches, configuring firewalls, and performing regular audits.
  • Conduct performance benchmarking, diagnostics, and capacity planning to keep systems up to date with evolving needs.
  • Perform other duties as required.
Supervision Received
  • This position reports to the Executive Director, AI Computing Resource (AICR)
Supervision Exercised
  • None
Employment Type
  • Full-Time, Hybrid (primarily remote with occasional on-site)
Qualifications & Skills
Required
  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent work experience).
  • Minimum 7 years relevant experience required.
  • Proven experience as an HPC System Engineer or similar role in high-performance computing environments
  • In-depth knowledge of HPC architectures, Linux/Unix operating systems, and cluster management tools.
  • Experience with parallel programming, MPI, job schedulers, and batch processing systems.
  • Strong knowledge of storage systems (e.g., Lustre, GPFS) and network file systems.
  • Experience with containerization and virtualization technologies (e.g., Docker, Kubernetes).
  • Strong troubleshooting skills
  • Excellent communication skills, with the ability to interact with technical and non-technical stakeholders.
Preferred
  • Experience with cloud-based HPC environments (e.g., AWS, Google Cloud).
  • Familiarity with GPU-based computing and relevant software (e.g., CUDA).
  • Knowledge of programming languages such as Python, Bash, or Perl.
  • Experience with monitoring and alerting tools (e.g., Nagios, Prometheus).
  • Experience with networking and network management tools
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Service Delivery Manager
Service Delivery Manager

The Massachusetts Green High Performance Computing Center Inc • Holyoke (MA), Northern (KY)

Hybrid
USD 145,000 - 196,000
Senior Full Stack Developer
Senior Full Stack Developer

The Massachusetts Green High Performance Computing Center Inc • Holyoke (MA), Northern (KY)

Hybrid
USD 120,000 - 163,000
Hybrid work model
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
Sr. HPC Systems Engineer (IT@JH Research Computing)
Sr. HPC Systems Engineer (IT@JH Research Computing)

Johns Hopkins University • Baltimore (MD)

On-site
USD 85,500 - 149,800
Senior HPC Systems Engineer — AI Compute Infra (Hybrid)
Senior HPC Systems Engineer — AI Compute Infra (Hybrid)

The Massachusetts Green High Performance Computing Center Inc • Holyoke (MA), Northern (KY)

Hybrid
USD 120,000 - 163,000
Lead Systems Engineer (HPC)
Lead Systems Engineer (HPC)

Princeton University • Princeton (NJ)

On-site
USD 135,000 - 150,000
Comprehensive benefits program
Senior HPC Systems Administrator
Senior HPC Systems Administrator

RedLine Performance Solutions, LLC. • Berkeley (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
paid time off
401k match
health care benefits
HPC Scientific Software Engineer (IT@JH Research Computing)
HPC Scientific Software Engineer (IT@JH Research Computing)

The Johns Hopkins University • Baltimore (MD)

On-site
USD 80,000 - 120,000
HPC Systems Administrator
HPC Systems Administrator

Scorpion Therapeutics • California (MO)

Hybrid
USD 150,000 - 190,000
Hybrid work model
HPC Engineer
HPC Engineer

Tata Consultancy Services • Indianapolis (IN)

On-site
USD 75,000 - 80,000
Discretionary annual incentive
Comprehensive medical coverage (M/D/V,
Family support leaves
+6