Senior HPC Systems Engineer — AI Compute Infra (Hybrid)

The Massachusetts Green High Performance Computing Center Inc

Holyoke, Northern (MA, KY)

Hybrid

USD 120,000 - 163,000

Full time

16 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

The Massachusetts Green High Performance Computing Center Inc. in Holyoke, MA is seeking an HPC Systems Engineer to lead deployment, maintenance, and optimization of HPC clusters, storage systems, and networking for AI/ML workloads.

This hands-on role collaborates with researchers and engineers to ensure secure, efficient, and reliable infrastructure. Responsibilities include installing, configuring hardware/software, monitoring performance, backing up data, and developing documentation.

Qualifications

  • Bachelor's degree in CS/Engineering or related field.
  • Minimum 7 years of relevant experience.
  • Proven HPC System Engineer experience in HPC environments.
  • Deep knowledge of HPC architectures, Linux/Unix, cluster management.
  • Experience with MPI, job schedulers, and batch systems.
  • Storage expertise (Lustre/GPFS) and network file systems.
  • Containerization/virtualization (Docker, Kubernetes).
  • Strong troubleshooting skills.
  • Excellent communication with technical and non-technical stakeholders.

Responsibilities

  • Manage infrastructure, support, maintenance, and operations of systems, hardware, and software.
  • Install, configure, and maintain HPC hardware and software, including clusters, storage systems, and job schedulers.
  • Monitor system performance, identify bottlenecks, and implement optimizations to improve efficiency and resource utilization.
  • Provide technical support to end-users, including troubleshooting hardware and software issues.
  • Perform system backups and disaster recovery procedures to ensure data integrity and availability.
  • Develop and maintain system documentation, including installation guides, configuration files, and standard operating procedures.
  • Assist in the design and deployment of scalable HPC environments to meet the needs of scientific and engineering applications.
  • Collaborate with teams to identify and implement software and hardware upgrades.
  • Ensure security of HPC systems by applying patches, configuring firewalls, and performing regular audits.
  • Conduct performance benchmarking, diagnostics, and capacity planning to keep systems up to date with evolving needs.
  • Perform other duties as required.

Skills

HPC architectures
Linux/Unix
Cluster management
MPI / parallel programming
Job schedulers
Storage systems
Docker / Kubernetes
Python / Bash / Perl
Networking tools
Troubleshooting
Communication skills
Cloud HPC (AWS/GCP)

Education

Bachelor's degree in Computer Science, Engineering, or related field

Tools

Nagios
Prometheus

Job description

The Massachusetts Green High Performance Computing Center Inc. in Holyoke, MA is seeking an HPC Systems Engineer to lead deployment, maintenance, and optimization of HPC clusters, storage systems, and networking for AI/ML workloads.

This hands-on role collaborates with researchers and engineers to ensure secure, efficient, and reliable infrastructure. Responsibilities include installing, configuring hardware/software, monitoring performance, backing up data, and developing documentation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Systems Engineer
Senior HPC Systems Engineer

The Massachusetts Green High Performance Computing Center Inc • Holyoke (MA), Northern (KY)

Hybrid
USD 120,000 - 163,000
HPC Operations & AI Service Delivery Lead
HPC Operations & AI Service Delivery Lead

Massachusetts Institute of Technology • Cambridge (MA)

On-site
USD 110,000 - 150,000
Senior Architect & Technical Lead, Portable HPC Integrations
Senior Architect & Technical Lead, Portable HPC Integrations

Massachusetts Institute of Technology • Cambridge (MA)

On-site
USD 120,000 - 160,000
HPC Service Delivery Lead for AI Compute
HPC Service Delivery Lead for AI Compute

The Massachusetts Green High Performance Computing Center Inc • Holyoke (MA), Northern (KY)

Hybrid
USD 145,000 - 196,000
Senior Full-Stack Engineer for AI Compute & Dashboards
Senior Full-Stack Engineer for AI Compute & Dashboards

The Massachusetts Green High Performance Computing Center Inc • Holyoke (MA), Northern (KY)

Hybrid
USD 120,000 - 163,000
Hybrid work model
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Senior HPC Architecture Lead — Hybrid, AI & Cloud
Senior HPC Architecture Lead — Hybrid, AI & Cloud

Memorial Sloan Kettering Cancer Center • New York (NY)

Hybrid
USD 136,000 - 225,000
HPC AI Infrastructure Architect
HPC AI Infrastructure Architect

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
Senior HPC Performance Engineer for AI & Big Compute
Senior HPC Performance Engineer for AI & Big Compute

Hewlett Packard Enterprise • Spring (TX)

On-site
USD 106,000 - 243,000