Site Reliability Engineer

Arlequin AI

Paris

Hybride

EUR 90 000 - 130 000

Plein temps

Il y a 4 jours
Soyez parmi les premiers à postuler
Générateur de candidature

N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Avantages offerts par ce poste

Equity stake
Flexible remote policy
Training, conference, and cert budget

Résumé du poste

Arlequin AI is building a cloud-native, sovereign platform for research and production. We seek a Senior SRE to ensure availability, performance, and scalability across our infra, while enabling teams with self-service tooling.

You will own reliability, observability, and automation in a small Platform team (3 today) with a culture of toil reduction, IaC, GitOps, and OpenTelemetry. Remote-friendly, Paris-based, hybrid/remote/on-site options.

Qualifications

  • 5+ years of experience in SRE / DevOps / Infrastructure
  • Advanced command of Kubernetes in production environments
  • Strong IaC experience (Terraform / OpenTofu, Terragrunt)
  • Hands-on GitOps (ArgoCD or Flux)
  • Observability culture (LGTM stack)
  • Scripting (Bash, Python) and Linux fundamentals
  • Autonomy and ownership over production scope, good communication

Responsabilités

  • Reliability & production: define, instrument, and track SLIs/SLOs/SLAs; improve resilience; reduce toil; handle incidents and post-mortems
  • Infrastructure & IaC: design/deploy/maintain infra on Scaleway; IaC with OpenTofu and Terragrunt; manage production Kubernetes
  • GitOps & delivery: operate CD with GitOps; standardize CI/CD pipelines; manage secrets securely
  • Observability: improve LGTM stack; design dashboards and alerts; promote OpenTelemetry
  • Culture & collaboration: share SRE/DevOps practices; document infrastructure and runbooks; contribute to platform roadmap
  • On-call: current policy; if rotation occurs, limited and paid

Connaissances

Kubernetes expertise
Linux fundamentals
Scripting: Bash, Python
DevOps culture
ownership and autonomy

Outils

Kubernetes
Terraform/OpenTofu
Terragrunt
ArgoCD
Flux
OpenTelemetry
LGTM stack (Loki, Grafana, Tempo, Mimir)

Description du poste

Ensure the availability, performance, and scalability of Arlequin's sovereign, cloud-native platform

5+ years, Senior

Full-time

Paris (hybrid / full-remote / on-site flexible)

About Arlequin

Arlequin AI is both a topological deep learning research lab and an AI platform. Our first product, HuDex, converts massive volumes of raw, unstructured, multilingual data into strategic decisions in minutes instead of days, providing an operational advantage for government agencies and businesses.
~30 people, post-seed, Series A in progress. On-site in Paris.

The role

We have chosen a cloud-native, sovereign, end-to-end automated infrastructure, hosted on Scaleway. Joining us means taking part in building and operating a reliable, secure, and observable platform serving our research, product, and development teams.

The Platform team owns the technical foundations Arlequin AI runs on: DevOps, DevSecOps, MLOps, DataOps, FinOps, security and compliance, compute, and IT. Our role is not to build infrastructure for its own sake — we build self-service tools so engineers can ship without waiting on us, a Forward Deployed Engineer can deploy a client without our help, C-levels understand the real cost of what they sell, and researchers don’t have to deal with engineering questions to run their experiments. Today we are a team of 3, and we are looking for people to help structure the team.

As a Senior SRE, you join the Platform team to ensure the availability, performance, and scalability of our platforms, while giving development teams the tooling to be autonomous. You work on both the run and the build side, with a genuine culture of toil reduction.

Your responsibilities

Reliability & production: define, instrument, and track SLIs / SLOs / SLAs; continuously improve resilience (capacity planning, load testing, chaos engineering, disaster recovery, eliminating SPOFs); reduce toil through automation and self-service; handle incidents and run blameless post-mortems; support the scaling of training, inference, and scientific computing workloads

Infrastructure & IaC: design, deploy, and maintain infrastructure on Scaleway; industrialize IaC with OpenTofu and Terragrunt; administer and evolve production Kubernetes clusters

GitOps & delivery: operate and evolve continuous deployment following GitOps principles with ArgoCD; standardize CI/CD pipelines and release workflows; manage configuration and secrets securely

Observability: improve and maintain our LGTM stack (Loki, Grafana, Tempo, Mimir); design dashboards and a relevant alerting policy with Alertmanager; promote OpenTelemetry observability and train engineering teams on it

Culture & collaboration: spread SRE / DevOps best practices across teams; document the infrastructure and runbooks; contribute to architecture decisions and the platform’s technical roadmap

On-call: no rotation today, incidents are handled during business hours. When a rotation becomes necessary, it will never exceed one week on-call out of five, it will be paid, and you’ll take part in designing it

Stack

Languages: Bash, Python

What we’re looking for

5+ years of experience in SRE / DevOps / Infrastructure, including significant experience with mission-critical production

Advanced command of Kubernetes in production environments

Solid experience with IaC (Terraform / OpenTofu, ideally Terragrunt)

Hands-on GitOps practice (ArgoCD or Flux)

Strong observability culture (ideally the LGTM stack)

Comfortable with scripting / automation (Bash, Python) and solid Linux fundamentals

Autonomy and a strong sense of ownership over the production scope, excellent communication, a feedback culture, technical curiosity, and pragmatism

Bonus

Hands-on experience with Scaleway

Knowledge of Cilium & service mesh

FinOps awareness

Process

First-fit interview (30 min)

Technical interview with the Platform team (coding + design)

Meeting with the hiring manager

Offer

Package

Competitive compensation including equity stake

Flexible remote policy: hybrid / full-remote / full on-site

Training, conference, and certification budget, with time dedicated to CNCF/LF open source contributions

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

HPC Engineer
HPC Engineer

Arlequin AI • Paris

Hybride
EUR 90 000 - 130 000
Senior SRE: Cloud-Native Infra, GitOps & Observability
Senior SRE: Cloud-Native Infra, GitOps & Observability

Arlequin AI • Paris

Hybride
EUR 90 000 - 130 000
Equity stake
Flexible remote policy
Training, conference, and cert budget
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alice & Bob • Paris

Hybride
EUR 85 000 - 120 000
BSPCE plan
Direct IP compensation bonuses
Flexible remote policy up to 40%
+5
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alice & Bob SAS • Paris

Sur place
EUR 90 000 - 130 000
Equity BSPCE
Significant patents bonuses
Remote policy up to 40%
+8
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Lille

Hybride
EUR 70 000 - 90 000
Up to 3 days remote per week
Office near public transport
Swile meal card
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Toulouse

Hybride
EUR 70 000 - 110 000
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Rennes

Hybride
EUR 65 000 - 90 000
Hybrid work
Free meals
Swile card
+3
Senior Product Manager
Senior Product Manager

Arlequin AI • Paris

Hybride
EUR 90 000 - 135 000
Equity stake
On-site in Paris
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Rouen

Hybride
EUR 70 000 - 100 000
Hybrid work
Lunch service
Swile lunch card
+2
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Lyon

Hybride
EUR 70 000 - 100 000
Hybrid work (up to 3 days remote)
Office near public transport
Healthy meals at HQ
+4