Artificial Intelligence Engineer

Penta Consulting

Dublin

On-site

EUR 90,000 - 130,000

Full time

16 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Penta Consulting is seeking an experienced HPC/AI Operations Specialist to provide onsite operational support and post-deployment stabilisation for a large-scale compute environment in Ireland.

You will join as a resident specialist, ensuring operational readiness, stability and performance while transitioning the platform to a sustainable steady-state operating model, with responsibilities spanning monitoring, maintenance, incident response and knowledge transfer to the customer operations team.

Qualifications

  • Extensive experience administering large-scale HPC environments.
  • Experience with cluster management technologies.
  • Hands-on experience with Slurm workload scheduler.
  • Experience with container orchestration platforms such as Kubernetes.
  • Knowledge of GPU infrastructure and AI/ML compute environments.
  • Experience with enterprise storage and networking within HPC.
  • Strong troubleshooting and operational support capabilities.
  • Experience with incident, problem and change management frameworks.
  • Excellent documentation and knowledge transfer skills.
  • Experience supporting research or AI-focused computing environments.

Responsibilities

  • Provide Day 2 operational support for the HPC and AI platform.
  • Monitor platform health, performance, availability and capacity.
  • Support planned maintenance activities, upgrades and change management processes.
  • Respond to, troubleshoot, and resolve incidents and operational issues.
  • Conduct root cause analysis and contribute to problem management.
  • Manage day-to-day cluster, compute, storage, and infrastructure operations.
  • Support workload scheduling and orchestration platforms including Slurm and Kubernetes.
  • Maintain operational documentation, procedures and runbooks.
  • Deliver knowledge transfer and operational readiness activities to the customer operations team.
  • Support the transition of the platform into a sustainable operational model.

Skills

HPC environments
Linux administration
Workload scheduling
Slurm
Kubernetes
GPU infrastructure
Enterprise storage & networking
Incident/problem/change mgmt
Documentation & knowledge transfer
Scripting (Bash, Python, Ansible)

Tools

Slurm
Kubernetes

Job description

Penta Consulting are a technology service provider and leading outsourced partner helping to deliver professional and managed solutions across EMEA.

We are seeking an experienced HPC / AI Operations Specialist to provide onsite operational support and post-deployment stabilisation for a leading-edge HPC and AI platform. Working as a resident specialist, you will play a critical role in ensuring the operational readiness, stability, and performance of a large-scale compute environment while supporting the transition into a sustainable steady-state operating model.

This role is ideal for a professional with deep expertise in HPC operations, Linux administration, workload scheduling, GPU-based infrastructure, and enterprise-scale cluster environments.

Key Responsibilities
  • Provide Day 2 operational support for the HPC and AI platform.
  • Monitor platform health, performance, availability, and capacity.
  • Support planned maintenance activities, upgrades, and change management processes.
  • Respond to, troubleshoot, and resolve incidents and operational issues.
  • Conduct root cause analysis and contribute to problem management activities.
  • Manage day-to-day cluster, compute, storage, and infrastructure operations.
  • Support workload scheduling and orchestration platforms including Slurm and Kubernetes.
  • Maintain operational documentation, procedures, and runbooks.
  • Deliver knowledge transfer and operational readiness activities to the customer operations team.
  • Support the transition of the platform into a sustainable operational model.
Required Skills & Experience
  • Extensive experience administering large-scale HPC environments.
  • Proven experience with cluster management technologies.
  • Hands‑on experience with workload schedulers such as Slurm.
  • Experience supporting container orchestration platforms such as Kubernetes.
  • Knowledge of GPU infrastructure and AI/ML compute environments.
  • Experience supporting enterprise storage and networking within HPC platforms.
  • Strong troubleshooting and operational support capabilities.
  • Experience working within structured incident, problem, and change management frameworks.
  • Excellent documentation and knowledge transfer skills.
  • Experience supporting research, scientific, or AI‑focused computing environments.
  • Familiarity with infrastructure monitoring and observability tools.
  • Experience with automation and scripting (Bash, Python, Ansible).
  • Exposure to data centre operations and environmental considerations.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Onsite HPC & AI Platform Engineer
Senior Onsite HPC & AI Platform Engineer

Penta Consulting • Dublin

On-site
EUR 90,000 - 130,000
Data Center Engineer (Cooling)
Data Center Engineer (Cooling)

Penta Consulting • Dublin

On-site
EUR 70,000 - 100,000
HPC & AI Data Center Cooling Operations Engineer
HPC & AI Data Center Cooling Operations Engineer

Penta Consulting • Dublin

On-site
EUR 70,000 - 100,000
Hardware Solutions Engineer
Hardware Solutions Engineer

AMAX • Galway

On-site
EUR 90,000 - 120,000
Hardware Solutions Engineer
Hardware Solutions Engineer

AMAX • Cork

On-site
EUR 85,000 - 120,000
Hardware Solutions Engineer
Hardware Solutions Engineer

AMAX • Limerick

On-site
EUR 90,000 - 120,000
HPC & AI Infrastructure Solutions Architect
HPC & AI Infrastructure Solutions Architect

AMAX • Galway

On-site
EUR 90,000 - 120,000
Senior HPC & AI Infrastructure Architect
Senior HPC & AI Infrastructure Architect

AMAX • Cork

On-site
EUR 85,000 - 120,000
Technical Delivery Director - AI, Agentic Systems & Production AI
Technical Delivery Director - AI, Agentic Systems & Production AI

EPAM Systems • Ireland

Hybrid
EUR 90,000 - 120,000
Technical Delivery Director - AI, Agentic Systems & Production AI
Technical Delivery Director - AI, Agentic Systems & Production AI

EPAM Systems • Dublin

Hybrid
EUR 180,000 - 240,000
Income protection
Life assurance
Health scheme
+3