SRE (DataPlatform)

Veepee

Paris

Hybrid

EUR 65,000 - 90,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health Insurance
Up to 2 days of remote work per week
E-learning platform access

Job summary

Veepee is looking for a Site Reliability Engineer (SRE) to join their Data Platform team in Paris. You will ensure the reliability, scalability, and operability of critical data services in a distributed environment. Responsibilities include implementing SRE best practices, automating infrastructure provisioning, and optimizing performance. Ideal candidates will have strong Kubernetes experience, SRE principles understanding, and collaboration skills. The position offers up to 2 days of remote work per week and health insurance.

Qualifications

  • Strong experience with Kubernetes in production environments.
  • Solid understanding of SRE principles, monitoring, alerting, SLAs/SLOs.
  • Experience with Infrastructure as Code using Terraform or similar.

Responsibilities

  • Ensure reliability and performance of data platform services.
  • Define and implement SRE best practices: SLIs/SLOs, error budgets, observability.
  • Automate infrastructure provisioning using Terraform and improve GitOps workflows.

Skills

Kubernetes in production environments
SRE principles (monitoring, alerting, SLAs/SLOs)
Infrastructure as Code (Terraform or similar)
Observability tools (Prometheus, Grafana)
Programming skills (Python, Java, or Go)
Experience with distributed data systems
Collaboration mindset and ability to work across teams

Tools

Prometheus
Grafana
Terraform
Trino
Kafka
Flink

Job description

Being an SRE at VeepeeTech means being part of a transversal SRE community while integrating a product-oriented Data Platform team.

You will contribute to the reliability, scalability, and operability of critical data services by applying SRE and DevOps practices, while sharing knowledge across teams.

The Data Platform is currently evolving toward a modern lakehouse architecture deployed on VeepeeCloud (our on-prem platform), based on technologies such as Trino, Iceberg, and object storage, with strong ambitions around performance, cost efficiency, and platform ownership.

You will work in a distributed environment (France & Spain), within a team of 40–50 data professionals across engineering, analytics, data science, and governance.

You will play a key role in ensuring the reliability and scalability of this next-generation data platform, while supporting the transition from public cloud to hybrid/on-prem architectures.

Platform Reliability & Operations
  • Ensure reliability and performance of our data platform services (Trino, Iceberg, S3, Kafka, Flink)
  • Define and implement SRE best practices: SLIs/SLOs, error budgets, observability
  • Build and maintain monitoring, alerting, and incident response frameworks (Prometheus, Grafana, etc.)
  • Contribute to the migration from public datawarehouse cloud to VeepeeCloud lakehouse stack
  • Support coexistence between cloud and on-prem systems and ensure consistency and reliability
  • Help design resilient architectures for ingestion, transformation, and serving layers
  • Operate and improve services running on Kubernetes (GKE/EKS & on-prem clusters)
  • Automate infrastructure provisioning using Terraform, Atlantis, and/or Crossplane
  • Improve GitOps workflows for platform deployment and configuration
FinOps & Performance Optimization
  • Collaborate with teams to optimize compute/storage usage (Trino queries, BigQuery slots, etc.)
  • Build tools and dashboards to track cost, usage, and efficiency
  • Support the transition toward cost-efficient on-prem workloads
  • Improveself-service capabilities for data teams (e.g., provisioning Trino/Iceberg resources)
  • Help teams adopt best practices in reliability, observability, and deployment
  • Write clear technical documentation and runbooks
Resilience & DRP
  • Contribute to Disaster Recovery Plan (DRP) definition and implementation
  • Ensure multi-DC resilience (FR1 / NL1) and data replication strategies
  • Participate in incident management and postmortems
  • Strong experience with Kubernetes in production environments
  • Experience with distributed data systems (or strong willingness to learn)
  • Solid understanding of SRE principles (monitoring, alerting, SLAs/SLOs)
  • Experience with Infrastructure as Code (Terraform or similar)
  • Familiarity with GitOps workflows
  • Experience with observability tools (Prometheus, Grafana, logging systems)
  • Comfortable working in cloud environments
  • Strong collaboration mindset and ability to work across teams
  • Experience with Trino, Iceberg, or data lakehouse architectures
  • Experience with Ceph S3 or object storage systems
  • Knowledge of Kafka / Flink / Airflow
  • Experience with FinOps practices and cost optimization
  • Experience with Crossplane or platform self-service models
  • Programming skills (Python, Java, or Go)
  • Experience with multi-region / multi-DC architectures
  • Dynamic and creative environment within international teams
  • The variety of self-education courses on our e-learning platform
  • The participation in meetups and conferences locally and internationally
  • Up to 2 days of remote work per week
  • Health Insurance
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE - DataPlatform
SRE - DataPlatform

Veepee • Paris

On-site
EUR 90,000 - 120,000
Flexible office – up to 2 days remote
International teams (France & Spain)
E-learning platform
+1
SRE
SRE

Veepee • Paris

On-site
EUR 50,000 - 70,000
Dynamic and creative environment
E-learning courses
Participation in meetups and conferences
+3
SRE - W/M/X
SRE - W/M/X

Lever, Inc. • Paris

On-site
EUR 65,000 - 90,000
Health Insurance
Up to 2 days remote work per week
Flexible working hours
+1
SRE Lead for Lakehouse Data Platform & Hybrid Infra
SRE Lead for Lakehouse Data Platform & Hybrid Infra

Veepee • Paris

Hybrid
EUR 90,000 - 120,000
Flexible office – up to 2 days remote
International teams (France & Spain)
E-learning platform
+1
SRE / DevOps Engineer - CDI - H/F/X
SRE / DevOps Engineer - CDI - H/F/X

Veepee • Paris

On-site
EUR 50,000 - 80,000
Health insurance
Self-education courses
Participation in meetups and conferences
+1
Hybrid SRE: Data Platform Lakehouse & Reliability
Hybrid SRE: Data Platform Lakehouse & Reliability

Veepee • Paris

Hybrid
EUR 65,000 - 90,000
Health Insurance
Up to 2 days of remote work per week
E-learning platform access
Senior SRE - BeReal
Senior SRE - BeReal

BeReal • Paris

On-site
EUR 90,000 - 130,000
Competitive salary based on experience
Lunch voucher (Swile)
Gymlib 100% covered by Voodoo
+2
Senior SRE - BeReal
Senior SRE - BeReal

United States Digital Space LLC • Paris

On-site
EUR 90,000 - 120,000
Lunch voucher
Gym membership
Premium health insurance
+1
Software Engineer Go, Rust or Scala (Infrastructure - Foundation) - Freelance (H/F/X)
Software Engineer Go, Rust or Scala (Infrastructure - Foundation) - Freelance (H/F/X)

Veepee • Paris

On-site
EUR 85,000 - 130,000
Variable bonus
International teams environment
E-learning courses
+2
Software Engineer Go, Rust or Scala (Infrastructure - Foundation) - Permanent (H/F/X)
Software Engineer Go, Rust or Scala (Infrastructure - Foundation) - Permanent (H/F/X)

Veepee • Paris

Hybrid
EUR 90,000 - 120,000
Flexible Office with up to 2 days at 0