Senior GPU Infra and Data Center Reliability Engineer

Coreweave

United States

On-site

USD 109,000 - 179,000

Full time

12 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Life Insurance
Voluntary Life Insurance
Disability Insurance
FSA
HSA
Tuition Reimbursement
ESPP
Mental Health Benefits
Family Formation Support
Parental Leave
Childcare Support
401(k) Matching
Flexible PTO
Catered Lunch

Job summary

CoreWeave is seeking a Senior Operations Engineer to join the MetalDev Operations team. You will own reliability, observability, and automation for frontend data-center operations, including hardware, firmware, and BMCs.

The role blends 80% production support with 20% process improvement and documentation. You'll work with NVIDIA GPUs, CDUs, NVLink switches, and standardized hardware, collaborating across Hardware Engineering and Fleet Ops.

Qualifications

  • 5+ years in cloud operations, SRE, or related field.
  • Working knowledge of Kubernetes and AWS or Google Cloud Platform.
  • Experience deploying and supporting containerized applications in Kubernetes.
  • Experience incident management: response, escalation, post-incident reviews.
  • Experience using Prometheus, Grafana, and PromQL for monitoring and troubleshooting.
  • Strong Linux administration and scripting skills.
  • Experience on-call rotation supporting production services.
  • Excellent written and verbal communication, especially during high-impact incidents.

Responsibilities

  • Triage and troubleshoot team-owned services, including init and reboot issues.
  • Monitor fleet health, identify unhealthy devices, and coordinate remediation.
  • Own monitoring dashboards, alerts, and KPIs using Prometheus and Grafana.
  • Investigate hardware, firmware, and issues with internal teams and vendors.
  • Create and maintain runbooks, escalation procedures, and service docs.
  • Provide incident communications and coordinate post-incident reviews.

Skills

Kubernetes
Public Cloud (AWS/GCP)
Containerization
Incident Management
Prometheus/Grafana
Linux Systems
On-Call Rotation

Tools

Prometheus
Grafana
PromQL
BMCs/Redfish

Job description

CoreWeave is seeking a Senior Operations Engineer to join the MetalDev Operations team. You will own reliability, observability, and automation for frontend data-center operations, including hardware, firmware, and BMCs.

The role blends 80% production support with 20% process improvement and documentation. You'll work with NVIDIA GPUs, CDUs, NVLink switches, and standardized hardware, collaborating across Hardware Engineering and Fleet Ops.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Center & GPU Infrastructure Engineer
Senior Data Center & GPU Infrastructure Engineer

CoreWeave • Livingston (NJ)

On-site
USD 109,000 - 179,000
Medical, dental, vision insurance
401(k) with match
Paid parental leave
+2
Senior GPU Compute Platform Engineer
Senior GPU Compute Platform Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 153,000 - 242,000
Senior Compute Architect - GPU Data Center (Go)
Senior Compute Architect - GPU Data Center (Go)

Coreweave • New York (NY), California (MO)

On-site
USD 153,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
Bare-Metal GPU Support Engineer for Data Center Ops
Bare-Metal GPU Support Engineer for Data Center Ops

Coreweave • United States

Remote
USD 110,000 - 170,000
Medical, dental, vision fully covered
Life insurance and disability coverage
401(k) with employer match
+5
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

AI Chopping Block • California (MO)

Hybrid
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with generous match
Flexible PTO
+2
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

CoreWeave • Bellevue (WA)

On-site
USD 99,000 - 132,000
Medical/Dental/Vision
401(k) Matching
Flexible PTO
+7
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+3
Senior Software Engineer, Network Observability for GPUs
Senior Software Engineer, Network Observability for GPUs

CoreWeave • Sunnyvale (CA)

On-site
USD 153,000 - 204,000
Medical, dental, vision insurance
Company-paid Life Insurance
ESPP
+5
GPU Bare‑Metal Support Engineer for AI Cloud
GPU Bare‑Metal Support Engineer for AI Cloud

CoreWeave • San Francisco (CA)

On-site
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+2
Senior Network Services Engineer (GPU Cloud)
Senior Network Services Engineer (GPU Cloud)

CoreWeave • New York (NY)

On-site
USD 182,000 - 242,000
Medical, dental, vision insurance
Equity and 401(k) matching
Flexible PTO