Fleet Reliability Engineer - HPC & GPU Clusters

Coreweave

Plano (TX)

On-site

USD 83,000 - 110,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) with employer match
Tuition Reimbursement
Flexible PTO
Catered lunch

Job summary

CoreWeave is hiring for a Fleet Reliability Operations role focused on provisioning, managing, and maintaining a fleet of server nodes and GPUs across data centers. You will troubleshoot complex hardware/software issues, drive rapid resolution, and contribute to performance monitoring and process improvements.

The role requires strong Linux administration, scripting skills, and experience with HPC environments.

Qualifications

  • Strong understanding of Linux system administration and internals.
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks.
  • Software development or scripting languages (bash, python, powershell, etc).
  • Bachelor’s degree or equivalent experience in related field is preferred.
  • Experience with Grafana/Prometheus observability tooling is a plus.
  • Kubernetes administration experience is preferred.

Responsibilities

  • Configure and maintain large-scale high-performance supercomputing clusters running GPUs.
  • Troubleshoot hardware and software issues; coordinate with data center, network, hardware and platform teams to drive resolution.
  • Monitor system performance and take remediation actions for cloud health.
  • Participate in on-call rotations, including after hours and weekends.
  • Create and maintain documentation of team processes and best practices.

Skills

Linux administration
Troubleshooting
Scripting (bash/python)

Education

Bachelor's degree or equivalent

Tools

Grafana
Prometheus
PromQL
Kubernetes

Job description

CoreWeave is hiring for a Fleet Reliability Operations role focused on provisioning, managing, and maintaining a fleet of server nodes and GPUs across data centers. You will troubleshoot complex hardware/software issues, drive rapid resolution, and contribute to performance monitoring and process improvements.

The role requires strong Linux administration, scripting skills, and experience with HPC environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Fleet Reliability Engineer – HPC GPU Clusters On-Call
Fleet Reliability Engineer – HPC GPU Clusters On-Call

CoreWeave • Washington

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with generous employer match
Flexible PTO
+1
Fleet Automation Engineer – HPC Infrastructure
Fleet Automation Engineer – HPC Infrastructure

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lunch stipend
Medical benefits (HDHP)
401(k) match
Hybrid GPU Fleet Operations Engineer
Hybrid GPU Fleet Operations Engineer

Crusoe • San Francisco (CA)

Hybrid
USD 215,000 - 260,000
Hybrid work schedule
Industry competitive pay
Restricted Stock Units
+2
Senior Field Engineer, HPC Compute Infrastructure
Senior Field Engineer, HPC Compute Infrastructure

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
401(k) match
Paid time off
+3
Fleet Automation Engineer – HPC & GPU Compute
Fleet Automation Engineer – HPC & GPU Compute

NorthMark Strategies • Dallas (TX)

On-site
USD 120,000 - 180,000
Company-Paid Lunch Stipend
Company-Paid Benefits: Medical, Dental
401(k) matching up to 6%
+2
Fleet Engineering Lead: Data Center & RMA Ops
Fleet Engineering Lead: Data Center & RMA Ops

CoreWeave • Dallas (TX)

On-site
USD 99,000 - 132,000
Medical, dental, vision insurance
Equity awards
Discretionary bonus
+1
Senior Field Engineer: GPU Compute Infra & HPC
Senior Field Engineer: GPU Compute Infra & HPC

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 275,000
401(k) with employer match
Medical, dental, and vision insurance
Tuition Reimbursement
+1
24/7 Fleet Reliability Manager
24/7 Fleet Reliability Manager

Coreweave • Town of Washington (NY)

On-site
USD 143,000 - 191,000
Medical Insurance
Life Insurance
Disability Insurance
+13
Senior Data Center & GPU Infrastructure Engineer
Senior Data Center & GPU Infrastructure Engineer

CoreWeave • Livingston (NJ)

On-site
USD 109,000 - 179,000
Medical, dental, vision insurance
401(k) with match
Paid parental leave
+2
Fleet Automation Engineer – HPC Infra
Fleet Automation Engineer – HPC Infra

NMC2 • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 150,000