Senior Ops Engineer - GPU Data Center Reliability

Coreweave

Livingston

On-site

GBP 90,000 - 140,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Coreweave is seeking a Senior Operations Engineer to join the MetalDev Operations team. You will own monitoring, incident response, and tooling, focusing on reliability and automation in a cutting-edge data center environment.

You will work with hardware and firmware vendors, collaborate with hardware engineering and fleet operations, and drive improvements in runbooks and on-call processes to reduce toil and incidents.

Qualifications

  • 5+ years in cloud operations, SRE, or infrastructure fields.
  • Experience with Prometheus, Grafana, and PromQL for monitoring.
  • Strong Linux system administration and scripting skills.

Responsibilities

  • Monitor fleet health and coordinate remediation across teams.
  • Own monitoring dashboards, alerts, and operational KPIs.
  • Perform root-cause analysis and post-incident reviews with documented actions.
  • Troubleshoot team-owned services and on-call incidents.
  • Create and maintain runbooks, escalation procedures, and docs.

Skills

Cloud operations
Site reliability engineering
Incident management
Monitoring (Prometheus/Grafana)
Linux administration
Kubernetes
Python or Go
Hardware management
On-call experience
Data center operations
Redfish/IPMI tooling

Education

Bachelor’s degree in CS/Engineering

Tools

Prometheus
Grafana
Kubernetes
BMC/Redfish tooling
IPMI
Linux tooling

Job description

Coreweave is seeking a Senior Operations Engineer to join the MetalDev Operations team. You will own monitoring, incident response, and tooling, focusing on reliability and automation in a cutting-edge data center environment.

You will work with hardware and firmware vendors, collaborate with hardware engineering and fleet operations, and drive improvements in runbooks and on-call processes to reduce toil and incidents.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, GPU Data Center Automation
Staff Software Engineer, GPU Data Center Automation

CoreWeave • York and North Yorkshire

On-site
GBP 120,000 - 180,000
Operations Engineer (MetalDev)
Operations Engineer (MetalDev)

Coreweave • Livingston

On-site
GBP 90,000 - 140,000
Senior Data & MLOps Engineer - AI Reliability Platform
Senior Data & MLOps Engineer - AI Reliability Platform

CoreWeave • Greater London

On-site
GBP 120,000 - 190,000
Medical Insurance
Dental Insurance
Pension
+5
GPU HPC Operations Manager - Reliability & Agile Leader
GPU HPC Operations Manager - Reliability & Agile Leader

Northern Data Group • Greater London

Hybrid
GBP 90,000 - 130,000
Staff Software Engineer (MetalDev)
Staff Software Engineer (MetalDev)

Coreweave • York and North Yorkshire

On-site
GBP 120,000 - 180,000
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability
Senior HPC Infra SRE: GPU Compute, 24/7 Reliability

Radiant • Greater London

On-site
GBP 90,000 - 140,000
Senior Researcher, GPU Infra & Reliability
Senior Researcher, GPU Infra & Reliability

Coreweaveu • Greater London

On-site
GBP 80,000 - 100,000
Collaborative work environment
Opportunities for growth
Support for independent thinking
Senior Software Engineer (Infrastructure Engineering)
Senior Software Engineer (Infrastructure Engineering)

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 130,000
Senior Infrastructure Engineer – SRE & Automation
Senior Infrastructure Engineer – SRE & Automation

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 130,000
Senior Data Center Engineer — GPU & Linux Networking
Senior Data Center Engineer — GPU & Linux Networking

Pursuu • Manchester

Hybrid
GBP 40,000 - 70,000
Company events
Company pension
Free parking
+2