SRE - DataPlatform

Veepee

Paris

Sur place

EUR 90 000 - 120 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Flexible office – up to 2 days remote
International teams (France & Spain)
E-learning platform
Meetups & conferences

Résumé du poste

Veepee is building a transversal SRE community within a product-oriented Data Platform team of 40–50 engineers, analysts, and data scientists across France and Spain. You will drive reliability and scalability of a Lakehouse platform based on Trino, Iceberg, and on-prem storage, while steering hybrid/on-prem migrations.

You will own core data services like Trino, Iceberg, and Kafka ecosystems, enforce SLIs/SLOs, and create robust observability with Prometheus and Grafana.

Qualifications

  • Production Kubernetes experience.
  • Familiarity with Kube-in-Kube technologies (vCluster or similar).
  • Strong SRE fundamentals including SLIs/SLOs and error budgets.
  • Experience with Prometheus and Grafana.
  • Experience with Infrastructure as Code (Terraform or similar).
  • Experience with Crossplane and GitOps workflows.
  • Experience with S3 and object storage technologies.
  • Experience with PostgreSQL and Patroni.
  • Experience with Kafka, Kafka Connect, and Schema Registry.
  • Fluent in English.

Responsabilités

  • Own reliability of core data services including Trino, Iceberg, and Kafka ecosystems.
  • Define and enforce SLIs/SLOs, error budgets, and on-call runbooks.
  • Build full-stack observability with Prometheus and Grafana.
  • Operate production Kubernetes clusters across GKE/EKS and on-prem.
  • Architect data platform resilience with multi-DC and DRP.
  • Lead cloud migration from public cloud to hybrid/on-prem.

Connaissances

Fluent English
SRE principles
GitOps familiarity

Outils

Kubernetes (production)
Kube-in-Kube
Crossplane
Prometheus
Grafana
Terraform
Kafka Connect
Schema Registry
Kafka
Trino
Iceberg
PostgreSQL
Patroni
S3/Ceph
On-prem storage
Airflow

Description du poste

Join a transversal SRE community embedded in a product-oriented Data Platform team of 40–50 engineers, analysts, and data scientists across France and Spain. You'll drive the reliability and scalability of a next-generation Lakehouse platform – anchored on Trino, Iceberg, and on-prem object storage – while leading the transition from public cloud to a resilient hybrid/on-prem architecture.

Platform Reliability & SRE foundations
  • Own reliability of core data services: Trino, Iceberg, S3 / Ceph, Kafka, Kafka Connect, Schema Registry
  • Define and enforce SLIs/SLOs, error budgets, and on-call runbooks – solid SRE foundations are non-negotiable
  • Build full-stack observability with Prometheus and Grafana: metrics, dashboards, alerting pipelines, and anomaly detection
  • Manage and harden PostgreSQL clusters via Patroni for high-availability control-plane services
Kafka ecosystem – Connect & Schema governance
  • Operate and scale Kafka Connect clusters: connector lifecycle, offset management, dead-letter queues, and task rebalancing
  • Maintain the Schema Registry as the single source of truth for Avro/Protobuf/JSON schemas – enforce compatibility rules and schema evolution policies
  • Monitor consumer lag, connector throughput, and broker health via Prometheus JMX exporters and Grafana dashboards
  • Ensure end-to-end data contract integrity between producers and Iceberg/S3 consumers
Kubernetes, Kube-in-Kube & Crossplane
  • Operate production Kubernetes clusters (GKE/EKS + on-prem) – capacity planning, upgrades, PodDisruptionBudgets, resource quotas
  • Architect and manage Kube-in-Kube topologies to provide strong tenant isolation for data platform workloads – each team gets a dedicated virtual cluster without the overhead of a full physical cluster
  • Automate infrastructure and resource provisioning with Crossplane: define composite resources (XRDs) so data teams can self-serve Kafka topics, Trino namespaces, and S3 buckets through Kubernetes-native APIs
  • Maintain GitOps pipelines for platform deployment and configuration drift detection
Lakehouse architecture & cloud migration
  • Migrate from public cloud data warehouse to VeepeeCloud Iceberg-based lakehouse – managing coexistence, schema evolution, and time-travel
  • Architect resilient ingestion, transformation, and serving layers around Trino + S3
  • Optimize Trino query performance: memory limits, spilling, cost-based optimizer tuning
Agentic & developer enablement
  • Build agentic self-service tooling so data teams can provision Trino/Iceberg resources and Kafka Connect pipelines autonomously via Crossplane – reducing toil and ops bottlenecks
  • Develop FinOps dashboards (compute, storage, query cost) with Grafana and Prometheus-based cost exporters
  • Write clear technical documentation, runbooks, and internal ADRs
Multi-DC resilience & DRP
  • Design and implement multi-datacenter strategies across FR1 / NL1 – active-active and active-passive topologies
  • Leverage Fast Erasure Coding on object storage (Ceph/S3) to maximize durability with minimal replication overhead
  • Ensure data replication consistency across sites for Iceberg table metadata, Trino catalogs, and Schema Registry subjects
  • Lead DRP exercises: failover playbooks, RTO/RPO validation, postmortems
Must have
  • Strong experience with Kubernetes in production environments
  • Experience with Kube-in-Kube technologies (vCluster or similar)
  • Solid understanding of SRE principles (SLIs/SLOs, error budgets)
  • Experience with Prometheus and Grafana
  • Experience with Infrastructure as Code (Terraform or similar)
  • Experience with Crossplane
  • Familiarity with GitOps workflows
  • Experience with S3 and object storage technologies
  • Experience with PostgreSQL and Patroni
  • Experience with Kafka, Kafka Connect, and Schema Registry
  • Fluent in English
Desired experience
  • Experience with multi-datacenter architectures (FR1/NL1)
  • Experience designing disaster recovery plans and failover playbooks
  • Experience with Fast Erasure Coding (Ceph/S3)
  • Experience with Trino, Iceberg, and Lakehouse technologies
  • Experience with Airflow
  • Experience building agentic self-service platforms
  • Knowledge of FinOps and cost optimization practices
  • Programming experience in Python, Java, or Go
Benefits
  • Variable bonus
  • E-learning platform (self-education courses)
  • Meetups & conferences (local and international)
  • Flexible office – up to 2 days remote
  • International teams (France & Spain)

For the service of diversity and inclusion, Veepee is committed to reviewing all applications received on an equal basis.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

SRE - DataPlatform
SRE - DataPlatform

Veepee • Paris

Hybride
EUR 50 000 - 75 000
Variable bonus
Dynamic and creative environment
E-learning courses
+2
SRE (DataPlatform)
SRE (DataPlatform)

Veepee • Paris

Hybride
EUR 65 000 - 90 000
Health Insurance
Up to 2 days of remote work per week
E-learning platform access
SRE Lead for Lakehouse Data Platform & Hybrid Infra
SRE Lead for Lakehouse Data Platform & Hybrid Infra

Veepee • Paris

Hybride
EUR 90 000 - 120 000
Flexible office – up to 2 days remote
International teams (France & Spain)
E-learning platform
+1
SRE
SRE

Veepee • Paris

Sur place
EUR 50 000 - 70 000
Dynamic and creative environment
E-learning courses
Participation in meetups and conferences
+3
Software Engineer Go, Rust or Scala (Infrastructure) - Foundation (H/F/X)
Software Engineer Go, Rust or Scala (Infrastructure) - Foundation (H/F/X)

Veepee • Paris

Sur place
EUR 90 000 - 130 000
Variable bonus
Flexible Office
International teams
+2
Software Engineer Go, Rust or Scala (Infrastructure - Foundation) - Freelance (H/F/X)
Software Engineer Go, Rust or Scala (Infrastructure - Foundation) - Freelance (H/F/X)

Veepee • Paris

Hybride
EUR 85 000 - 130 000
Variable bonus
International teams environment
E-learning courses
+2
SRE / DevOps Engineer - CDI - H/F/X
SRE / DevOps Engineer - CDI - H/F/X

Veepee • Paris

Hybride
EUR 50 000 - 80 000
Health insurance
Self-education courses
Participation in meetups and conferences
+1
Senior SRE - BeReal
Senior SRE - BeReal

BeReal. • Paris

Sur place
EUR 90 000 - 140 000
Swile Lunch voucher
Gymlib
Premium healthcare coverage
+1
Ingénieur Senior Base de données & Backend F/H - Senior Database & Backend Engineer (Infrastructure & Abstraction)
Ingénieur Senior Base de données & Backend F/H - Senior Database & Backend Engineer (Infrastructure & Abstraction)

GE Vernova • Montpellier

Sur place
EUR 70 000 - 110 000
Alternance - Data Engineer H/F/X
Alternance - Data Engineer H/F/X

Veepee • Saint-Denis

Sur place
EUR 28 000 - 36 000
Télétravail jusqu’à 2 jours/semaine
Plateforme d’apprentissage des langues
CSE et avantages
+2