HPC Fleet Reliability Engineer

CoreWeave

Livingston (NJ)

On-site

USD 83,000 - 110,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Life Insurance
Disability insurance
FSA/HSA
Tuition Reimbursement
Stock Purchase Program (ESPP)
Mental Wellness
Parental Leave
Flexible PTO
Casual work environment

Job summary

CoreWeave is actively seeking engineers for the Fleet Reliability Operations team to manage and optimize CoreWeave's fleet of server nodes. You will contribute to provisioning, validation, and on-call support to ensure fast, reliable deployment of GPU-accelerated clusters.

Ideal candidates bring Linux expertise, scripting ability, and experience with data center or HPC environments. Competitive salary, strong benefits, and a collaborative, fast-paced culture are offered.

Qualifications

  • Strong understanding of Linux system administration and internals.
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably.
  • Software development or scripting languages (bash, python, powershell, etc).

Responsibilities

  • Configure and maintain large-scale high-performance supercomputing clusters running state-of-the-art GPUs.
  • Troubleshoot hardware and software issues; escalate with data center, network, hardware and platform teams to drive resolution.
  • Monitor and analyze system performance and take remediation actions for cloud health.
  • Approach work with flexibility and optimism amid shifting priorities.
  • Create and maintain documentation of team processes, knowledge and best practices for system management.
  • Think critically about day-to-day work and collaborate to improve team processes and efficiency.
  • Participate in on-call rotations which include after hours and weekend work.

Skills

Linux administration
Troubleshooting
Scripting

Education

Bachelor's degree in related field

Tools

Kubernetes administration

Job description

CoreWeave is actively seeking engineers for the Fleet Reliability Operations team to manage and optimize CoreWeave's fleet of server nodes. You will contribute to provisioning, validation, and on-call support to ensure fast, reliable deployment of GPU-accelerated clusters.

Ideal candidates bring Linux expertise, scripting ability, and experience with data center or HPC environments. Competitive salary, strong benefits, and a collaborative, fast-paced culture are offered.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Fleet Reliability Engineer for HPC GPU Clusters & Live Ops
Fleet Reliability Engineer for HPC GPU Clusters & Live Ops

CoreWeave • Washington

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+7
Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Washington

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+7
Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Plano (TX)

Hybrid
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with a generous employer match
Flexible PTO
+2
Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Livingston (NJ)

On-site
USD 83,000 - 110,000
Medical insurance
Life Insurance
Disability insurance
+7
Senior Field Engineer - HPC Compute Infrastructure
Senior Field Engineer - HPC Compute Infrastructure

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
Operations Engineering Manager, Fleet Reliability
Operations Engineering Manager, Fleet Reliability

CoreWeave • New York (NY)

On-site
USD 150,000 - 200,000
Engineering Manager, Fleet Engineering
Engineering Manager, Fleet Engineering

CoreWeave • Livingston (NJ)

On-site
USD 165,000 - 210,000
Staff Software Engineer, Scalable GPU Infrastructure
Staff Software Engineer, Scalable GPU Infrastructure

CoreWeave • New York (NY)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
Operations Engineering Manager, Fleet Reliability
Operations Engineering Manager, Fleet Reliability

CoreWeave • Livingston (NJ)

On-site
USD 143,000 - 191,000
Medical, dental, and vision insurance
Equity awards
401(k) with employer match
+4
HPC Performance Engineer
HPC Performance Engineer

CoreWeave • Bellevue (WA)

Hybrid
USD 165,000 - 242,000
Medical, dental, and vision insurance
Flexible PTO
Catered lunch each day
+2