Senior Data Center Ops Engineer - Hardware Reliability

Marimo Inc.

New York, Livingston (NY, NJ)

On-site

USD 109,000 - 179,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical/Dental/Vision
Company-paid Life Insurance
401(k) with company match
Flexible PTO
Equity awards

Job summary

CoreWeave is seeking a Senior Operations Engineer to join the MetalDev Operations team. You’ll own monitoring, incident response, and runbooks, collaborating with hardware and firmware teams to improve tooling and processes.

You will work with GPU servers, power shelves, BMCs, and other data-center components, building scalable observability and remediation capabilities to reduce toil and improve service reliability.

Qualifications

  • 5+ years in cloud operations, SRE, or infra ops.
  • Hands-on with Kubernetes and a public cloud (AWS or GCP).
  • Experience deploying containerized apps in Kubernetes.
  • Incident management, response, and post-incident reviews.
  • Prometheus, Grafana, and PromQL for monitoring.
  • Strong Linux administration and scripting skills.
  • Ability to troubleshoot complex cross-stack issues.
  • Experience participating in an on-call rotation.
  • Excellent communication and documentation.

Responsibilities

  • Triage and troubleshoot team-owned services in production.
  • Own observability, dashboards, alerts, and KPIs for services.
  • Investigate hardware, firmware, and vendor issues.
  • Collaborate with hardware/firmware teams to validate fixes.
  • Create and maintain runbooks, playbooks, and incident docs.

Skills

Cloud operations
SRE practices
Incident management
On-call rotation
Observability
Automation
Linux administration

Education

Bachelor's degree in CS/Engineering or equivalent

Tools

Kubernetes
AWS
GCP
Prometheus
Grafana
PromQL
Linux
Python
Golang
BMCs/Redfish

Job description

CoreWeave is seeking a Senior Operations Engineer to join the MetalDev Operations team. You’ll own monitoring, incident response, and runbooks, collaborating with hardware and firmware teams to improve tooling and processes.

You will work with GPU servers, power shelves, BMCs, and other data-center components, building scalable observability and remediation capabilities to reduce toil and improve service reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Center & GPU Infrastructure Engineer
Senior Data Center & GPU Infrastructure Engineer

CoreWeave • Livingston (NJ)

On-site
USD 109,000 - 179,000
Medical, dental, vision insurance
401(k) with match
Paid parental leave
+2
Senior Infrastructure Engineer — AI Cloud Reliability
Senior Infrastructure Engineer — AI Cloud Reliability

Socket.dev • New York (NY)

On-site
USD 182,000 - 242,000
Medical/Dental/Vision
Life Insurance
Tuition Reimbursement
+5
Senior Hardware Engineer: Scalable Server Automation
Senior Hardware Engineer: Scalable Server Automation

CoreWeave • Bellevue (WA)

On-site
USD 109,000 - 204,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous match
+3
Senior GPU Compute Platform Engineer
Senior GPU Compute Platform Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 153,000 - 242,000
GPU Data Center Operations Engineer
GPU Data Center Operations Engineer

CoreWeave • New York (NY)

On-site
USD 109,000 - 145,000
Medical insurance
401(k) with employer match
Flexible PTO
+1
Senior Data Center Operational Readiness Engineer
Senior Data Center Operational Readiness Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 198,000 - 264,000
Medical, dental, and vision insurance
401(k) with employer match
Paid Parental Leave
+2
Onsite Data Center Technician - Linux & Hardware
Onsite Data Center Technician - Linux & Hardware

CoreWeave • Cheyenne (WY)

On-site
USD 60,000 - 90,000
Data Center Technician - Hands-on Hardware & Networking
Data Center Technician - Hands-on Hardware & Networking

CoreWeave • Town of Afton (NY)

On-site
USD 65,000 - 83,000
Medical insurance
Life Insurance
Disability insurance
+4
Senior Compute Architect - GPU Data Center (Go)
Senior Compute Architect - GPU Data Center (Go)

Coreweave • New York (NY), California (MO)

On-site
USD 153,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
Lead, Technical Support Eng — Bare Metal & GPU Infra
Lead, Technical Support Eng — Bare Metal & GPU Infra

Socket.dev • Sunnyvale (CA), Seattle (WA), San Francisco (CA)

On-site
USD 157,000 - 210,000
Medical, dental, and vision insurance
401(k) with generous employer match
Flexible PTO
+2