HPC Fleet Reliability Engineer

CoreWeave

Plano (TX)

On-site

USD 83,000 - 110,000

Full time

24 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with generous match
Tuition Reimbursement
ESPP
Paid Parental Leave
Flexible PTO
Catered lunch

Job summary

CoreWeave is seeking a Fleet Reliability Operations engineer to manage the day-to-day provisioning, upgrades and uptime of our expanding fleet of server nodes. You will configure GPU clusters, troubleshoot issues and coordinate with data center teams to drive rapid resolution while maintaining cloud health.

Ideal candidates will have Linux expertise, scripting ability, and experience with data-center infrastructure.

Qualifications

  • Strong understanding of Linux system administration and internals.
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably.
  • Software development or scripting languages (bash, python, powershell, etc).

Responsibilities

  • Configure and maintain large-scale high-performance supercomputing clusters running state-of-the-art GPUs.
  • Troubleshoot hardware and software issues; coordinate with data center, network, hardware and platform teams to drive resolution.
  • Monitor and analyze system performance and take remediation actions for cloud health.
  • Participate in oncall rotations which include after hours and weekend work.
  • Create and maintain documentation of team processes, knowledge and best practices for system management.
  • Think critically about daily work and collaborate to improve team processes and efficiency.

Skills

Linux administration
Troubleshooting
Scripting

Education

Bachelor’s degree or equivalent

Tools

Grafana
Prometheus

Job description

CoreWeave is seeking a Fleet Reliability Operations engineer to manage the day-to-day provisioning, upgrades and uptime of our expanding fleet of server nodes. You will configure GPU clusters, troubleshoot issues and coordinate with data center teams to drive rapid resolution while maintaining cloud health.

Ideal candidates will have Linux expertise, scripting ability, and experience with data-center infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Fleet Reliability Engineer
HPC Fleet Reliability Engineer

CoreWeave • Livingston (NJ)

On-site
USD 83,000 - 110,000
Medical insurance
Life Insurance
Disability insurance
+7
Fleet Reliability Engineer – HPC GPU Clusters On-Call
Fleet Reliability Engineer – HPC GPU Clusters On-Call

CoreWeave • Washington

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with generous employer match
Flexible PTO
+1
GPU HPC Systems Engineer — Fleet Reliability & Automation
GPU HPC Systems Engineer — Fleet Reliability & Automation

OpenAI • California (MO)

On-site
USD 180,000 - 260,000
Fleet Automation Engineer – HPC Infrastructure
Fleet Automation Engineer – HPC Infrastructure

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lunch stipend
Medical benefits (HDHP)
401(k) match
Fleet Reliability Ops Manager for Scale & Automation
Fleet Reliability Ops Manager for Scale & Automation

CoreWeave • Washington

On-site
USD 143,000 - 191,000
Medical/Dental/Vision
Life Insurance
Disability Insurance
+5
Fleet Hardware Reliability Engineer (Automation & HPC)
Fleet Hardware Reliability Engineer (Automation & HPC)

OpenAI • California (MO)

On-site
USD 180,000 - 270,000
Fleet Engineering Lead: Data Center & RMA Ops
Fleet Engineering Lead: Data Center & RMA Ops

CoreWeave • Dallas (TX)

On-site
USD 99,000 - 132,000
Medical, dental, vision insurance
Equity awards
Discretionary bonus
+1
HPC Operations Pro: Fleet Health & User Support
HPC Operations Pro: Fleet Health & User Support

Career Techniques • New York (NY)

Hybrid
USD 175,000 - 225,000
Lead, 24/7 Fleet Reliability & Automation
Lead, 24/7 Fleet Reliability & Automation

CoreWeave • Sunnyvale (CA)

On-site
USD 143,000 - 191,000
Medical, dental, and vision insurance
Life Insurance
Tuition Reimbursement
+4
HPC Performance Engineer: Kernel & Systems
HPC Performance Engineer: Kernel & Systems

CoreWeave • New York (NY)

On-site
USD 165,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Tuition Reimbursement
+2