Fleet Reliability Engineer — HPC & GPU Clusters

CoreWeave

Dublin

On-site

EUR 80,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
Life Assurance at 4x Salary
Critical Illness Cover
Employee Assistance Programme
Tuition Reimbursement
Work culture focused on innovative Dis

Job summary

CoreWeave is seeking a curious, persistent problem solver to join the Fleet Reliability Operations team in Ireland for provisioning, management and uptime of our GPU-rich server clusters. You will help push nodes through provisioning, validation and rapid troubleshooting as part of a focused engineering group.

You will work with a team dedicated to high-availability infrastructure, collaborating closely with data center, network and platform teams, while documenting processes and improving

Qualifications

  • Strong understanding of Linux system administration and internals.
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably.
  • Software development or scripting languages (bash, python, powershell, etc).
  • Bachelor’s degree in a related field or equivalent experience.

Responsibilities

  • Configure and maintain large-scale high-performance supercomputing clusters running state-of-the-art GPUs.
  • Troubleshoot hardware and software issues; escalate and coordinate with data center, network, hardware and platform teams to drive resolution.
  • Monitor and analyze system performance and take appropriate remediation actions for cloud health.
  • Create and maintain documentation of team processes, knowledge and best practices for system management.
  • Think critically about day-to-day work and collaborate to improve team processes and efficiency.

Skills

Linux administration
Troubleshooting
Scripting

Education

Bachelor’s degree or equivalent in a related field

Tools

Grafana
Prometheus
PromSQL
Kubernetes

Job description

CoreWeave is seeking a curious, persistent problem solver to join the Fleet Reliability Operations team in Ireland for provisioning, management and uptime of our GPU-rich server clusters. You will help push nodes through provisioning, validation and rapid troubleshooting as part of a focused engineering group.

You will work with a team dedicated to high-availability infrastructure, collaborating closely with data center, network and platform teams, while documenting processes and improving

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infrastructure Engineer: Fleet Automation & Diagnostics
GPU Infrastructure Engineer: Fleet Automation & Diagnostics

Crusoe • Dublin

On-site
EUR 65,000 - 85,000
Private health and dental insurance
Pension contributions
Income protection
+1
Senior Software Engineer - GPU Cloud Networking
Senior Software Engineer - GPU Cloud Networking

CoreWeave • Dublin

On-site
EUR 120,000 - 180,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+6
Hybrid Senior GPU Cloud Infra Engineer (Dublin)
Hybrid Senior GPU Cloud Infra Engineer (Dublin)

Uniting Holding • Dublin

Hybrid
EUR 75,000 - 95,000
25 days paid annual leave
Free inference tokens
Remote flexibility
HPC & AI Infrastructure Solutions Architect
HPC & AI Infrastructure Solutions Architect

AMAX • Galway

On-site
EUR 90,000 - 120,000
HPC & AI Data Center Cooling Operations Engineer
HPC & AI Data Center Cooling Operations Engineer

Penta Consulting • Dublin

On-site
EUR 70,000 - 100,000
Senior Infrastructure Engineer (GPU Cloud)
Senior Infrastructure Engineer (GPU Cloud)

Uniting Holding • Dublin

Hybrid
EUR 75,000 - 95,000
25 days paid annual leave
Free inference tokens
Remote flexibility
Linux Compute Farm Engineer
Linux Compute Farm Engineer

Software Placements • Cork

On-site
EUR 92,000 - 157,000
Competitive daily rate
Senior Network Production Operations Engineer — Global HPC Edge
Senior Network Production Operations Engineer — Global HPC Edge

Crusoe • Dublin

On-site
EUR 120,000 - 190,000
Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Dublin

On-site
EUR 80,000 - 120,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+5
Senior Network Production Engineer - Global HPC Edge
Senior Network Production Engineer - Global HPC Edge

ProducePay • Dublin

On-site
EUR 120,000 - 180,000
Pension contributions
Private health insurance
Dental insurance
+2