GPU Data Center Operations Engineer

CoreWeave

New York (NY, NJ)

On-site

USD 109,000 - 145,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical insurance
401(k) with employer match
Flexible PTO
Catered lunch in office

Job summary

CoreWeave is seeking an Operations Engineer for the MetalDev team to support data center bring-ups and production services. You will troubleshoot generation issues, improve observability, and help automate reliability tasks, collaborating with Fleet Operations and Hardware Engineering.

You will work with NVIDIA GPUs and related hardware, focusing 80% on production operations and 20% on monitoring improvements, with a base salary ranging from $109,000 to $145,000 plus bonuses and benefits.

Qualifications

  • Two+ years in technical support, systems administration, cloud operations or similar.
  • Proficient Linux administration, logs, networking, and CLI troubleshooting.
  • Kubernetes concepts and cloud platforms experience essential.
  • Experience with monitoring/observability tools like Prometheus and Grafana.

Responsibilities

  • Monitor services and fleet health; coordinate remediation.
  • Triage routine production issues and incident response.
  • Investigate problems across Linux, networks, hardware and software.
  • Improve dashboards, runbooks, and automation with guidance.
  • Contribute to root-cause analyses and post-incident reviews.

Skills

Linux administration
Kubernetes
Cloud platforms
Monitoring & observability
Scripting (bash)
On-call / incident response
Troubleshooting infrastructure
Documentation & communication
Team collaboration
Root-cause analysis

Education

Bachelor's degree in CS / Engineering / IT or related

Tools

Prometheus
Grafana
PromQL

Job description

CoreWeave is seeking an Operations Engineer for the MetalDev team to support data center bring-ups and production services. You will troubleshoot generation issues, improve observability, and help automate reliability tasks, collaborating with Fleet Operations and Hardware Engineering.

You will work with NVIDIA GPUs and related hardware, focusing 80% on production operations and 20% on monitoring improvements, with a base salary ranging from $109,000 to $145,000 plus bonuses and benefits.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Center & GPU Infrastructure Engineer
Senior Data Center & GPU Infrastructure Engineer

CoreWeave • Livingston (NJ)

On-site
USD 109,000 - 179,000
Medical, dental, vision insurance
401(k) with match
Paid parental leave
+2
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

AI Chopping Block • California (MO)

Hybrid
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with generous match
Flexible PTO
+2
Senior Data Center Ops Engineer - Hardware Reliability
Senior Data Center Ops Engineer - Hardware Reliability

Marimo Inc. • New York (NY), Livingston (NJ)

On-site
USD 109,000 - 179,000
Medical/Dental/Vision
Company-paid Life Insurance
401(k) with company match
+2
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+3
GPU Bare‑Metal Support Engineer for AI Cloud
GPU Bare‑Metal Support Engineer for AI Cloud

CoreWeave • San Francisco (CA)

On-site
USD 99,000 - 132,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+2
Senior Compute Architect - GPU Data Center (Go)
Senior Compute Architect - GPU Data Center (Go)

Coreweave • New York (NY), California (MO)

On-site
USD 153,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
GPU Bare-Metal Support Engineer
GPU Bare-Metal Support Engineer

CoreWeave • Bellevue (WA)

On-site
USD 99,000 - 132,000
Medical/Dental/Vision
401(k) Matching
Flexible PTO
+7
Senior GPU Compute Platform Engineer
Senior GPU Compute Platform Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 153,000 - 242,000
GPU HPC Cluster Ops Engineer — Equity & Growth
GPU HPC Cluster Ops Engineer — Equity & Growth

CoreWeave • Livingston (NJ)

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
Life Insurance
Flexible Spending Account
+8
Cloud GPU Support Engineer
Cloud GPU Support Engineer

Emploive • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Medical, dental, and vision insurance
Tuition Reimbursement
401(k) with employer match