Service Delivery Manager

The Massachusetts Green High Performance Computing Center Inc

Holyoke, Northern (MA, KY)

Hybrid

USD 145,000 - 196,000

Full time

21 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

The Massachusetts Green High Performance Computing Center Inc is seeking an experienced Service Delivery Manager to oversee the operational management and continuous improvement of the AICR. This role coordinates compute and data services to a diverse user base, partnering with AI Hub staff and stakeholders to ensure reliable, secure delivery.

You will lead service management, monitor performance, manage vendors, and collaborate with stakeholders to align technical solutions with organizational

Qualifications

  • Bachelor’s degree in CS/Engineering/IT or equivalent.
  • Minimum 7 years relevant experience in service delivery.
  • Experience managing IT infrastructure or HPC clusters.
  • Strong leadership and stakeholder communication.
  • Knowledge of HPC cluster management, job scheduling, performance tuning.

Responsibilities

  • Oversee end-to-end service delivery of the HPC cluster.
  • Serve as primary liaison with AICR stakeholders and communicate performance.
  • Lead resolution of service disruptions and incidents.
  • Analyze delivery performance and drive process improvements.
  • Plan capacity and allocate resources for growth.
  • Assist with budgeting and cost forecasting.
  • Manage external vendors and SLAs.
  • Identify risks and implement mitigation strategies.
  • Perform other duties as required.

Skills

HPC technologies
Leadership
Service management
Communication
Problem solving
Vendor management
GPU HPC experience

Education

Bachelor’s degree in CS/Engineering/IT or equivalent experience

Tools

Nagios
Prometheus
Zabbix
AWS
Azure
Google Cloud

Job description

Ensure high-quality, reliable operation of the state-of-the-art HPC environment behind the AI Computing Resource (AICR) that supports the Massachusetts AI Hub. This critical role will coordinate delivery of compute and data services to a diverse user base, including working in close partnership with facilitation and science support staff from AI Hub organizations, liaising with academic and industry stakeholders, overseeing performance monitoring and continuous improvement, and contributing to strategic planning and vendor management. This is an exciting opportunity to drive service excellence in support of world-class AI research and innovation.

Pay Range

$144,505 - $195,900

Position Overview

We are looking for an experienced and dynamic Service Delivery Manager to oversee the operational management and continuous improvement of the AICR. The Service Delivery Manager will be responsible for ensuring the smooth, efficient, and reliable delivery of resources and services to internal and external stakeholders. This role combines technical expertise with leadership skills to ensure that technical solutions align with organizational needs and provide an exceptional user experience. As a Service Delivery Manager for AICR, you will be the primary point of contact for ensuring that service delivery meets operational targets. You will coordinate with AICR technical staff and collaborate with users to ensure services are delivered effectively and that issues are addressed promptly.

Principal Responsibilities
  • Service Management: Oversee the end-to-end service delivery of the HPC cluster, ensuring the systems are running efficiently and meeting performance, availability, and security requirements.
  • Client and Stakeholder Liaison: Act as the primary point of contact to AICR stakeholders, for issues related to service delivery, ensuring their needs and expectations are understood and met. Regularly communicate service performance, enhancements, and upcoming updates.
  • Problem Management: Lead the identification and resolution of service disruptions, performance issues, or other incidents affecting AICR. Coordinate with relevant parties to restore services and prevent recurrence.
  • Continuous Improvement: Analyze service delivery performance and identify areas for improvement. Implement process improvements to enhance the efficiency, reliability, and scalability of the HPC cluster.
  • Capacity Planning and Resource Allocation: Work with stakeholders to understand future service demands and capacity requirements. Plan for scaling the HPC environment to meet growth needs, ensuring optimal resource allocation.
  • Budget and Cost Management: Work with the Executive Director to budget for HPC services, ensuring cost-effective delivery without compromising service quality. Work with finance and other stakeholders to forecast and track expenses.
  • Vendor Management: Collaborate with external vendors and service providers to ensure the HPC cluster is supported by the necessary hardware, software, and maintenance services. Manage vendor relationships and ensure the service levels are maintained.
  • Risk Management: Identify potential risks in the service delivery pipeline and develop mitigation strategies. Ensure that any risks related to performance, security, or downtime are appropriately addressed.
  • Perform other duties as required.
Supervision Received
  • This position reports to the Executive Director, AI Computing Resource (AICR)
Supervision Exercised
  • None
Employment Type
  • Full-Time, Hybrid (primarily remote with occasional on-site)
Qualifications & Skills
Required
  • Education:Bachelor’s degree in Computer Science, Engineering, IT, or a related field (or equivalent experience).
  • Experience:
  • Minimum 7 years relevant experience required.
  • Proven experience managing the service delivery of complex IT infrastructure or computing environments, preferably in HPC clusters.
  • Demonstrated leadership in a service management or delivery management role, managing both people and projects.
  • Strong understanding of HPC technologies, including cluster management, job scheduling, parallel computing, and performance tuning.
  • Skills:
  • Strong understanding of service management best practices
  • Excellent communication and interpersonal skills, with the ability to engage with technical and non-technical stakeholders.
  • Problem-solving skills
  • Ability to drive continuous improvement
Preferred
  • Experience delivering AI-specific cluster services
  • Experience with cloud-based HPC solutions or hybrid environments (e.g., AWS, Azure, Google Cloud).
  • Familiarity with monitoring tools and metrics, such as Nagios, Prometheus, or Zabbix.
  • Experience with service management frameworks (e.g. ITIL)
  • Experience in GPU-based computing, storage management, or network configuration for HPC clusters.
  • Ability to lead teams effectively
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Systems Engineer
Senior HPC Systems Engineer

The Massachusetts Green High Performance Computing Center Inc • Holyoke (MA), Northern (KY)

Hybrid
USD 120,000 - 163,000
Senior Full Stack Developer
Senior Full Stack Developer

The Massachusetts Green High Performance Computing Center Inc • Holyoke (MA), Northern (KY)

Hybrid
USD 120,000 - 163,000
Hybrid work model
HPC Service Delivery Lead for AI Compute
HPC Service Delivery Lead for AI Compute

The Massachusetts Green High Performance Computing Center Inc • Holyoke (MA), Northern (KY)

Hybrid
USD 145,000 - 196,000
Service Delivery Manager
Service Delivery Manager

Massachusetts Institute of Technology • Cambridge (MA)

On-site
USD 110,000 - 150,000
Lead Systems Engineer (HPC)
Lead Systems Engineer (HPC)

Princeton University • Princeton (NJ)

On-site
USD 135,000 - 150,000
Comprehensive benefits program
HPC AI Technologist
HPC AI Technologist

Cambridge Computer Services, Inc • Waltham (MA)

On-site
USD 90,000 - 210,000
Competitive salary
Multiple health insurance options
401(k) savings plan with employer matching
+2
HPC AI Technologist
HPC AI Technologist

Cambridge Computer Services, Inc • San Francisco (CA)

On-site
USD 90,000 - 210,000
Competitive salary
Multiple health insurance options
401(k) savings plan with employer matching
+2
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
HPC Scientific Software Engineer (IT@JH Research Computing)
HPC Scientific Software Engineer (IT@JH Research Computing)

The Johns Hopkins University • Baltimore (MD)

On-site
USD 80,000 - 120,000
HPC Infrastructure and Cluster Engineer
HPC Infrastructure and Cluster Engineer

Arena Technical Resources, LLC (ATR) • Springfield (VA)

On-site
USD 180,000 - 200,000