HPC GPU Fleet Reliability Engineer

CoreWeave Europe

Dublin

On-site

EUR 70,000 - 100,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
Life Assurance at 4x Salary
Critical Illness Cover
Employee Assistance Programme
Tuition Reimbursement

Job summary

CoreWeave is seeking engineers for the Fleet Reliability Operations team to manage provisioning, validation and uptime of its growing fleet of server nodes. You will help deploy nodes rapidly, troubleshoot issues, and document processes while collaborating with data center, network and platform teams.

You will contribute to high-performance GPU clusters, improve operational efficiency, and participate in on-call rotations, including after-hours.

Qualifications

  • Strong understanding of Linux system administration and internals.
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably.
  • Software development or scripting languages (bash, python, powershell, etc)

Responsibilities

  • Configure and maintain large-scale high-performance supercomputing clusters running state-of-the-art GPUs.
  • Troubleshoot hardware and software issues; coordinate with data center, network, hardware and platform teams to drive resolution.
  • Monitor and analyze system performance and take remediation actions for cloud health.

Skills

Linux administration
Troubleshooting hardware/ software
Scripting (bash, Python, PowerShell)

Education

Bachelor’s degree in related field or equivalent experience

Tools

Kubernetes
Grafana
Prometheus
PromQL

Job description

CoreWeave is seeking engineers for the Fleet Reliability Operations team to manage provisioning, validation and uptime of its growing fleet of server nodes. You will help deploy nodes rapidly, troubleshoot issues, and document processes while collaborating with data center, network and platform teams.

You will contribute to high-performance GPU clusters, improve operational efficiency, and participate in on-call rotations, including after-hours.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Fleet Reliability Engineer — HPC & GPU Clusters
Fleet Reliability Engineer — HPC & GPU Clusters

CoreWeave • Dublin

On-site
EUR 80,000 - 120,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+5
Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave Europe • Dublin

On-site
EUR 70,000 - 100,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+4
Senior GPU Cloud Reliability Engineer
Senior GPU Cloud Reliability Engineer

Crusoe Energy Systems • Ireland

On-site
EUR 90,000 - 130,000
Health benefits and wellness
Paid time off
401(k) match
+1
Senior Network Production Engineer - Global HPC Edge
Senior Network Production Engineer - Global HPC Edge

ProducePay • Dublin

On-site
EUR 120,000 - 180,000
Pension contributions
Private health insurance
Dental insurance
+2
Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Dublin

On-site
EUR 80,000 - 120,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+5
Staff Network Production Engineer – Global HPC Edge
Staff Network Production Engineer – Global HPC Edge

Crusoe • Dublin

On-site
EUR 120,000 - 180,000
Pension contributions
Private health and dental insurance
Income protection
+1
GPU-Accelerated AI/HPC Infrastructure Engineer
GPU-Accelerated AI/HPC Infrastructure Engineer

AMAX • Shannon

On-site
EUR 70,000 - 110,000
Senior Cloud HPC Engineer — 24/7 On-Call
Senior Cloud HPC Engineer — 24/7 On-Call

Crusoe • Dublin

On-site
EUR 60,000 - 90,000
Pension contributions
Private health insurance
Dental insurance
+2
Senior Cloud Support Engineer | 24/7 On-Call HPC/AI Infra
Senior Cloud Support Engineer | 24/7 On-Call HPC/AI Infra

Crusoe Energy Systems LLC • Dublin

On-site
EUR 70,000 - 105,000
Remote Technical Lead for GPU Infrastructure & Slurm
Remote Technical Lead for GPU Infrastructure & Slurm

Tether • Dublin

On-site
EUR 120,000 - 180,000
Remote work