Senior Data & MLOps Engineer - GPU Reliability Platform

Coreweaveu

Greater London

On-site

GBP 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Pension Contribution
Life Insurance
Critical Illness Cover
Employee Assistance
Tuition Reimbursement
Disruption-focused culture

Job summary

CoreWeave is seeking a Senior Data & MLOps Engineer to design and scale the GPU Intelligence Platform infrastructure. You will build data pipelines, feature processing, and production-grade ML model deployment across a fleet, with a focus on scalability, real-time and offline processing, and resource management.

You will work with Kubernetes-based microservices, collaborate with Platform teams, and implement monitoring and diagnostic signals to continuously improve system performance and

Qualifications

  • 7+ years of experience in data engineering, distributed systems, MLOps, or infrastructure ML roles in production environments.
  • Proven experience building high-throughput streaming or telemetry pipelines (e.g., Kafka, Pulsar, Kinesis, or equivalent).
  • Strong experience designing time-series feature pipelines and operating large-scale observability systems.
  • Experience building and maintaining feature stores and ensuring offline/online feature parity.
  • Hands-on experience deploying ML models to production, including versioning, monitoring, rollback, and drift detection.
  • Experience designing scalable microservices deployed in Kubernetes-based environments.
  • Strong proficiency in Python and at least one systems language (Go, Rust, or C++).
  • Experience working with distributed compute or training systems (e.g., NCCL, PyTorch Distributed, Spark, Ray, Slurm).
  • Familiarity with GPU telemetry systems such as NVML or DCGM and hardware-level monitoring concepts.
  • Demonstrated experience scaling systems from Proof-of-Concept to production-grade, fleet-level deployments.

Responsibilities

  • Design and implement scalable data ingestion pipelines.
  • Build feature processing and baseline computation systems.
  • Productionize models for prediction and detection.
  • Develop and operate low-latency service and robust offline workflows.
  • Architect horizontally scalable services with clear separation between components, leveraging orchestration for distribution.
  • Implement monitoring and feedback loops for continuous model and signal improvement.
  • Collaborate with Platform teams to integrate operational signals into monitoring and diagnostics.
  • Implement a scalable solution for mitigation and structured analysis.

Skills

Data engineering
Distributed systems
MLOps
Kubernetes
Python
Go/Rust/C++
Streaming pipelines
Feature stores
Model deployment
Observability

Tools

Kafka
Pulsar
Kinesis
NCCL
PyTorch
Spark
Ray
Slurm
NVML

Job description

CoreWeave is seeking a Senior Data & MLOps Engineer to design and scale the GPU Intelligence Platform infrastructure. You will build data pipelines, feature processing, and production-grade ML model deployment across a fleet, with a focus on scalability, real-time and offline processing, and resource management.

You will work with Kubernetes-based microservices, collaborate with Platform teams, and implement monitoring and diagnostic signals to continuously improve system performance and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Researcher, GPU Infra & Reliability
Senior Researcher, GPU Infra & Reliability

Coreweaveu • Greater London

On-site
GBP 80,000 - 100,000
Collaborative work environment
Opportunities for growth
Support for independent thinking
ML Ops Engineer: AI Platform & GPU Infra
ML Ops Engineer: AI Platform & GPU Infra

Anaplan • Greater London

On-site
GBP 90,000 - 150,000
MLOps & Infrastructure Engineer for Scalable AI Compute
MLOps & Infrastructure Engineer for Scalable AI Compute

Graphcore • West of England

On-site
GBP 70,000 - 110,000
Private medical insurance
Dental plan
Pension (matched up to 5%)
+4
Senior Data & MLOps Engineer
Senior Data & MLOps Engineer

Coreweaveu • Greater London

On-site
GBP 120,000 - 180,000
Medical Insurance
Dental Insurance
Pension Contribution
+5
MLOps & Infrastructure Engineer for Scalable AI Compute
MLOps & Infrastructure Engineer for Scalable AI Compute

graphcore • United Kingdom

Hybrid
GBP 70,000 - 100,000
Flexible working
Private medical insurance
Pension (matched up to 5%)
+6
ML Performance Engineer – Scale GPU/CPU Workloads
ML Performance Engineer – Scale GPU/CPU Workloads

Barlowe LLP • Greater London

On-site
GBP 90,000 - 150,000
Lunch provided
35 days’ annual leave
9% company pension contributions
+4
Staff Software Engineer - Network Automation & Cloud Infra
Staff Software Engineer - Network Automation & Cloud Infra

CoreWeave • Greater London

On-site
GBP 116,000 - 155,000
Family‑level Medical Insurance
Family‑level Dental Insurance
Generous Pension Contribution
+5
ML Infra & MLOps Engineer for AI Compute
ML Infra & MLOps Engineer for AI Compute

graphcore • Bristol

Hybrid
GBP 70,000 - 110,000
Flexible working
Private medical insurance
Pension (matched)
+1
Senior Researcher: AI Infrastructure & Reliability
Senior Researcher: AI Infrastructure & Reliability

CoreWeave • Greater London

On-site
GBP 70,000 - 90,000
Family-level Medical Insurance
Generous Pension Contribution
Tuition Reimbursement
Senior AI Platform Engineer - GPUaaS & MLOps Lead
Senior AI Platform Engineer - GPUaaS & MLOps Lead

Sharon AI, Inc • United Kingdom

On-site
GBP 90,000 - 150,000