GPU HPC Cluster Ops Engineer — Equity & Growth

CoreWeave

Livingston (NJ)

On-site

USD 83,000 - 110,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
Life Insurance
Flexible Spending Account
Health Savings Account
Tuition Reimbursement
Employee Stock Purchase Program (ESPP)
Mental Wellness Benefits
Paid Parental Leave
401(k) with employer match
Flexible PTO
Catered lunches

Job summary

CoreWeave is seeking a Fleet Reliability Operations professional to manage and optimize the company’s high-performance GPU clusters. You will handle provisioning, troubleshooting, and uptime while coordinating with data center, network, and platform teams to maximize node delivery to customers.

The role emphasizes hands-on cluster management, documentation, and collaboration in a fast-paced environment with on-call responsibilities. Bachelor’s degree or equivalent experience is expected.

Qualifications

  • Strong understanding of Linux system administration and internals.
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks reliably.
  • Scripting or software development experience (bash, Python, PowerShell).

Responsibilities

  • Configure and maintain large-scale high-performance clusters with GPUs.
  • Troubleshoot hardware and software issues and coordinate with data center and network teams for resolution.
  • Monitor system performance and take remediation actions for cloud health.
  • Adapt to shifting priorities with flexibility and optimism.
  • Create and maintain documentation of processes and best practices.
  • Collaborate to improve team processes and efficiency.
  • Participate in on-call rotations including after hours and weekends.

Skills

Linux administration
Troubleshooting
Scripting (bash, Python)

Education

Bachelor’s degree or equivalent in related field

Tools

Kubernetes
Grafana
Prometheus
PromSQL

Job description

CoreWeave is seeking a Fleet Reliability Operations professional to manage and optimize the company’s high-performance GPU clusters. You will handle provisioning, troubleshooting, and uptime while coordinating with data center, network, and platform teams to maximize node delivery to customers.

The role emphasizes hands-on cluster management, documentation, and collaboration in a fast-paced environment with on-call responsibilities. Bachelor’s degree or equivalent experience is expected.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Fleet Reliability Engineer - HPC & GPU Clusters
Fleet Reliability Engineer - HPC & GPU Clusters

Coreweave • Plano (TX)

On-site
USD 83,000 - 110,000
Health insurance
401(k) with employer match
Tuition Reimbursement
+2
Fleet Reliability Engineer – HPC GPU Clusters On-Call
Fleet Reliability Engineer – HPC GPU Clusters On-Call

CoreWeave • Washington

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with generous employer match
Flexible PTO
+1
Hybrid GPU Fleet Operations Engineer
Hybrid GPU Fleet Operations Engineer

Crusoe • San Francisco (CA)

Hybrid
USD 215,000 - 260,000
Hybrid work schedule
Industry competitive pay
Restricted Stock Units
+2
GPU Data Center Operations Engineer
GPU Data Center Operations Engineer

CoreWeave • New York (NY)

On-site
USD 109,000 - 145,000
Medical insurance
401(k) with employer match
Flexible PTO
+1
Senior Data Center & GPU Infrastructure Engineer
Senior Data Center & GPU Infrastructure Engineer

CoreWeave • Livingston (NJ)

On-site
USD 109,000 - 179,000
Medical, dental, vision insurance
401(k) with match
Paid parental leave
+2
Cloud GPU Support Engineer
Cloud GPU Support Engineer

Emploive • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Medical, dental, and vision insurance
Tuition Reimbursement
401(k) with employer match
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Senior Field Engineer: GPU Compute Infra & HPC
Senior Field Engineer: GPU Compute Infra & HPC

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 275,000
401(k) with employer match
Medical, dental, and vision insurance
Tuition Reimbursement
+1
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

AI Chopping Block • California (MO)

Hybrid
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with generous match
Flexible PTO
+2
Staff Network Platform Engineer – GPU Cloud Automation
Staff Network Platform Engineer – GPU Cloud Automation

CoreWeave • New York (NY)

On-site
USD 180,000 - 260,000
Medical insurance
Dental insurance
Vision insurance
+13