Fleet Reliability Engineer – HPC GPU Clusters On-Call

CoreWeave

Washington (District of Columbia)

On-site

USD 83,000 - 110,000

Full time

38 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with generous employer match
Flexible PTO
Catered lunch

Job summary

CoreWeave, The Essential Cloud for AI, is seeking a Fleet Reliability Operations engineer in Washington, DC to manage provisioning, updates, and uptime of our GPU-rich HPC clusters. You’ll troubleshoot issues, collaborate with data center, network, and platform teams, and help maximize node delivery to customers.

Join a fast-paced team with on-call rotations and a focus on documenting processes. Base salary ranges from $83k to $110k, with discretionary bonus, equity and comprehensive benefits.

Qualifications

  • Strong understanding of Linux system administration and internals.
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably.
  • Scripting or programming in Bash, Python, PowerShell, etc.

Responsibilities

  • Configure and maintain large-scale high-performance supercomputing clusters running GPUs.
  • Troubleshoot hardware and software issues; collaborate with data center, network, hardware and platform teams to resolve.
  • Monitor and analyze system performance and take remediation actions for cloud health.

Skills

Linux administration
Troubleshooting
Scripting (bash, Python)

Education

Bachelor’s degree or equivalent experience

Tools

Grafana
Prometheus
PromSQL
Kubernetes

Job description

CoreWeave, The Essential Cloud for AI, is seeking a Fleet Reliability Operations engineer in Washington, DC to manage provisioning, updates, and uptime of our GPU-rich HPC clusters. You’ll troubleshoot issues, collaborate with data center, network, and platform teams, and help maximize node delivery to customers.

Join a fast-paced team with on-call rotations and a focus on documenting processes. Base salary ranges from $83k to $110k, with discretionary bonus, equity and comprehensive benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Fleet Reliability Engineer
HPC Fleet Reliability Engineer

CoreWeave • Livingston (NJ)

On-site
USD 83,000 - 110,000
Medical insurance
Life Insurance
Disability insurance
+7
HPC Fleet Reliability Engineer
HPC Fleet Reliability Engineer

CoreWeave • Plano (TX)

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with generous match
Tuition Reimbursement
+4
GPU HPC Systems Engineer — Fleet Reliability & Automation
GPU HPC Systems Engineer — Fleet Reliability & Automation

OpenAI • California (MO)

On-site
USD 180,000 - 260,000
Lead GPU Cloud Fleet Engineering
Lead GPU Cloud Fleet Engineering

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health, dental, and vision coverage
401k with 2% company match (USA)
Flexible paid time off
Hybrid GPU Fleet Operations Engineer
Hybrid GPU Fleet Operations Engineer

Crusoe • San Francisco (CA)

Hybrid
USD 215,000 - 260,000
Hybrid work schedule
Industry competitive pay
Restricted Stock Units
+2
Fleet Hardware Reliability Engineer (Automation & HPC)
Fleet Hardware Reliability Engineer (Automation & HPC)

OpenAI • California (MO)

On-site
USD 180,000 - 270,000
Senior GPU HPC Platform Reliability Engineer
Senior GPU HPC Platform Reliability Engineer

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
Engineering Manager, AI Cloud GPU Fleet
Engineering Manager, AI Cloud GPU Fleet

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Fleet Automation Engineer – HPC Infrastructure
Fleet Automation Engineer – HPC Infrastructure

NorthMark Strategies LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Lunch stipend
Medical benefits (HDHP)
401(k) match
GPU Compute Engineer — Fleet Reliability & Automation
GPU Compute Engineer — Fleet Reliability & Automation

Fluidstack • San Francisco (CA)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Retirement or pension plan
Generous PTO policy