Fleet Reliability Engineer for HPC GPU Clusters & Live Ops

CoreWeave

Washington (District of Columbia)

On-site

USD 83,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
HSA/Health Savings Account
Tuition Reimbursement
Employee Stock Purchase Program
Mental Wellness Benefit
Parental Leave and childcare support
401(k) with employer match
Flexible PTO

Job summary

CoreWeave is seeking a Fleet Reliability Operations engineer to manage provisioning, maintenance, and uptime of its expanding fleet of server nodes. You’ll work on configuration, updates, and remote troubleshooting of top-tier HPC clusters and their networking, delivery platforms and tool dependencies.

You will join a focused team to deploy nodes quickly, rack them, and power them on while collaborating with data center, hardware, and platform teams to resolve issues and improve processes.

Qualifications

  • Strong understanding of Linux system administration and internals.
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably.
  • Experience with scripting languages (bash, python, powershell, etc.)

Responsibilities

  • Configure and maintain large‑scale high‑performance supercomputing clusters running state‑of‑the‑art GPUs.
  • Troubleshoot hardware and software issues; escalation and coordination with data center, network, hardware and platform teams to drive resolution.
  • Monitor and analyze system performance and take appropriate remediation actions for cloud health.
  • Approach your work with flexibility and optimism anticipating shifting business and technical priorities.
  • Create and maintain documentation of team processes, knowledge and best practices for system management.
  • Think critically about your day‑to‑day work and collaborate to improve team processes and efficiency.
  • Participate in on‑call rotations, including after‑hours and weekend work.

Skills

Linux administration
Troubleshooting
Scripting (bash, Python, PowerShell)

Education

Bachelor’s degree in a related field or equivalent experience

Tools

Grafana
Prometheus
promsql
Kubernetes

Job description

CoreWeave is seeking a Fleet Reliability Operations engineer to manage provisioning, maintenance, and uptime of its expanding fleet of server nodes. You’ll work on configuration, updates, and remote troubleshooting of top-tier HPC clusters and their networking, delivery platforms and tool dependencies.

You will join a focused team to deploy nodes quickly, rack them, and power them on while collaborating with data center, hardware, and platform teams to resolve issues and improve processes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Washington

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+7
Fleet Reliability Engineer — HPC & GPU Clusters
Fleet Reliability Engineer — HPC & GPU Clusters

CoreWeave • Plano (TX)

Hybrid
USD 83,000 - 110,000
Fleet Automation Engineer – HPC Infrastructure
Fleet Automation Engineer – HPC Infrastructure

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lunch stipend
Medical benefits (HDHP)
401(k) match
Head of Fleet Reliability & Automation
Head of Fleet Reliability & Automation

CoreWeave • Sunnyvale (CA)

On-site
USD 180,000 - 230,000
GPU HPC Fleet Reliability Engineer
GPU HPC Fleet Reliability Engineer

CoreWeave • Bellevue (WA)

Hybrid
USD 83,000 - 110,000
Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Plano (TX)

Hybrid
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with a generous employer match
Flexible PTO
+2
Fleet Reliability Operations Manager (24/7, AI Cloud)
Fleet Reliability Operations Manager (24/7, AI Cloud)

CoreWeave • Livingston (NJ)

On-site
USD 143,000 - 191,000
Medical, dental, and vision insurance
Equity awards
401(k) with employer match
+4
Fleet Reliability Engineer — Hybrid/Remote
Fleet Reliability Engineer — Hybrid/Remote

CoreWeave • Sunnyvale (CA)

Hybrid
USD 83,000 - 110,000
Operations Engineering Manager, Fleet Reliability
Operations Engineering Manager, Fleet Reliability

CoreWeave • New York (NY)

On-site
USD 150,000 - 200,000
Fleet Automation Engineer for HPC Infrastructure
Fleet Automation Engineer for HPC Infrastructure

NorthMark Compute and Cloud LLC • Fort Worth (TX), Town of Texas (WI)

On-site
USD 120,000 - 170,000
Lunch stipend
Medical benefits (HDHP)
Dental & Vision
+3