Senior Data & MLOps Engineer - AI Reliability Platform

CoreWeave

Greater London

On-site

GBP 120,000 - 190,000

Full time

12 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Pension
Life Assurance 4x Salary
Critical Illness Cover
Employee Assistance Programme
Tuition Reimbursement
Innovative disruption culture

Job summary

CoreWeave is recruiting a Senior Data & MLOps Engineer to design and scale the GPU Intelligence Platform infrastructure, building pipelines for data, features, and model training while delivering insights for system health and optimization.

You will transition prototypes to production across a fleet, focusing on scalable distributed services, separating real-time from offline processing, and dynamic resource management based on load and data frequency.

Qualifications

  • 7+ years of experience in data engineering, distributed systems, MLOps, or infrastructure ML roles in production environments.
  • Experience building high-throughput streaming or telemetry pipelines (e.g., Kafka, Pulsar, Kinesis, or equivalent).
  • Strong experience designing time-series feature pipelines and operating large-scale observability systems.
  • Experience building and maintaining feature stores and ensuring offline/online feature parity.
  • Hands‑on experience deploying ML models to production, including versioning, monitoring, rollback, and drift detection.
  • Experience designing scalable microservices deployed in Kubernetes-based environments.
  • Strong proficiency in Python and at least one systems language (Go, Rust, or C++).
  • Experience working with distributed compute or training systems (e.g., NCCL, PyTorch Distributed, Spark, Ray, Slurm).
  • Familiarity with GPU telemetry systems such as NVML or DCGM and hardware‑level monitoring concepts.
  • Demonstrated experience scaling systems from Proof‑of‑Concept to production‑grade, fleet‑level deployments.

Responsibilities

  • Design and implement scalable data ingestion pipelines.
  • Build feature processing and baseline computation systems.
  • Productionize models for prediction and detection.
  • Develop and operate low-latency service and robust offline workflows.
  • Architect horizontally scalable services with clear separation between components, leveraging orchestration for distribution.
  • Implement monitoring and feedback loops for continuous model and signal improvement.
  • Collaborate with Platform teams to integrate operational signals into monitoring and diagnostics.
  • Implement a scalable solution for mitigation and structured analysis.

Skills

Data engineering
Distributed systems
MLOps
Python
Go
Rust
C++
Kubernetes
Streaming pipelines
Observability

Tools

Kafka
Pulsar
Kinesis

Job description

CoreWeave is recruiting a Senior Data & MLOps Engineer to design and scale the GPU Intelligence Platform infrastructure, building pipelines for data, features, and model training while delivering insights for system health and optimization.

You will transition prototypes to production across a fleet, focusing on scalable distributed services, separating real-time from offline processing, and dynamic resource management based on load and data frequency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data & MLOps Engineer - GPU Reliability Platform
Senior Data & MLOps Engineer - GPU Reliability Platform

Coreweaveu • Greater London

On-site
GBP 120,000 - 180,000
Medical Insurance
Dental Insurance
Pension Contribution
+5
Production Data Scientist, AI Infra & Optimization
Production Data Scientist, AI Infra & Optimization

CoreWeave • Greater London

On-site
GBP 90,000 - 170,000
Family‑level Medical Insurance
Family‑level Dental Insurance
Generous Pension Contribution
+5
Senior Data Infrastructure Engineer, AI Compute Platform
Senior Data Infrastructure Engineer, AI Compute Platform

Mistral • Greater London

Hybrid
GBP 76,000 - 111,000
Healthcare coverage
Relocation support
Wellness programs
ML Ops Engineer: AI Platform & GPU Infra
ML Ops Engineer: AI Platform & GPU Infra

Anaplan • Greater London

On-site
GBP 90,000 - 150,000
Senior Researcher: AI Infrastructure & Reliability
Senior Researcher: AI Infrastructure & Reliability

CoreWeave • Greater London

On-site
GBP 70,000 - 90,000
Family-level Medical Insurance
Generous Pension Contribution
Tuition Reimbursement
Senior Software Engineer – Distributed AI Platform
Senior Software Engineer – Distributed AI Platform

CoreWeave • Greater London

On-site
GBP 100,000 - 160,000
Senior Data Platform Engineer — AI-Powered Systems
Senior Data Platform Engineer — AI-Powered Systems

Scale AI • Greater London

On-site
GBP 110,000 - 170,000
Staff Software Engineer, Physical AI Platform
Staff Software Engineer, Physical AI Platform

CoreWeave • Greater London

On-site
GBP 120,000 - 180,000
Senior Software Engineer — Physical AI Platform
Senior Software Engineer — Physical AI Platform

Jackalope Digital LLC • Greater London

Hybrid
GBP 98,000 - 130,000
Medical insurance
Life Insurance
Disability insurance
+12
MLOps Engineer: Production ML Pipelines & Observability
MLOps Engineer: Production ML Pipelines & Observability

CoreWeave • Greater London

On-site
GBP 100,000 - 150,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+4