GPU Cloud SRE Engineering Lead

Webhosting

Paris

Hybrid

EUR 90,000 - 130,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hybrid work
Dining service
Swile card

Job summary

Scaleway in Paris is seeking a senior Site Reliability Engineer to lead the GPU Cloud SRE team and build a production-grade infra powering our sovereign cloud. You will scale automated solutions across GPU clusters and ensure high availability.

You will collaborate with software engineering and product teams, mentor engineers, and contribute to the reliability roadmap. The role offers hybrid work with flexible remote days and offices near public transport.

Qualifications

  • Experience leading a 6-person SRE team.
  • Strong background in production-grade GPU infrastructure.
  • Proficiency with observability, logging and monitoring at scale.
  • Ability to design and implement automated server lifecycle solutions.

Responsibilities

  • Lead and manage a team of 6 Site Reliability Engineers.
  • Design automated solutions for server lifecycle management across GPU clusters.
  • Design and implement observability, logging, and monitoring solutions for large-scale GPU clusters.
  • Plan, prioritize, and manage the technical development roadmap for the SRE team.
  • Collaborate with software engineering, product, and cross-functional teams across Scaleway.
  • Handle recruitment and career management for team members.
  • Maintain, scale, and optimize high-availability production systems under heavy load.
  • Participate in on-call rotations to ensure production reliability and fast incident resolution.

Skills

Team leadership
SRE
GPU
Automation

Tools

Kubernetes
Monitoring stacks
Automation tools

Job description

Scaleway in Paris is seeking a senior Site Reliability Engineer to lead the GPU Cloud SRE team and build a production-grade infra powering our sovereign cloud. You will scale automated solutions across GPU clusters and ensure high availability.

You will collaborate with software engineering and product teams, mentor engineers, and contribute to the reliability roadmap. The role offers hybrid work with flexible remote days and offices near public transport.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Engineering Manager – GPU Cloud
SRE Engineering Manager – GPU Cloud

Webhosting • Paris

On-site
EUR 90,000 - 130,000
Hybrid work
Dining service
Swile card
Head of GPU Cloud Engineering
Head of GPU Cloud Engineering

Scaleway • Paris

Hybrid
EUR 180,000 - 240,000
GPU Cloud Operations Lead — Hybrid, Global Impact
GPU Cloud Operations Lead — Hybrid, Global Impact

Webhosting • Paris

Hybrid
EUR 110,000 - 150,000
Hybrid work
Modern offices
Healthy meals
+3
AI GPU SRE for Scalable Hybrid Infrastructure
AI GPU SRE for Scalable Hybrid Infrastructure

Scaleway • Paris

Hybrid
EUR 50,000 - 70,000
Hybrid work: up to 3 remote days per week
Chef-served meals
Access to gym and daycare
Head of Operations GPU Cloud
Head of Operations GPU Cloud

Webhosting • Paris

On-site
EUR 110,000 - 150,000
Hybrid work
Modern offices
Healthy meals
+3
Senior SRE Leader — Hybrid Cloud Platform Reliability
Senior SRE Leader — Hybrid Cloud Platform Reliability

Scaleway • Paris

Hybrid
EUR 70,000 - 90,000
Hybrid work: up to 3 days remote
Healthy meal service
Access to gym
+1
Site Reliability Engineer (SRE) - AI GPU Clusters
Site Reliability Engineer (SRE) - AI GPU Clusters

Scaleway • Paris

On-site
EUR 50,000 - 70,000
Hybrid work: up to 3 remote days per week
Chef-served meals
Access to gym and daycare
Head of Engineering - GPU Cloud
Head of Engineering - GPU Cloud

Scaleway • Paris

On-site
EUR 180,000 - 240,000
Site Reliability Engineer: Cloud Infra & Automation
Site Reliability Engineer: Cloud Infra & Automation

Scaleway • Lyon

Hybrid
EUR 75,000 - 110,000
Hybrid work (up to 3 days remote)
Lunch card (Swile)
Gym access
+1
Site Reliability Engineer — Cloud Stability & Automation
Site Reliability Engineer — Cloud Stability & Automation

Scaleway • Bordeaux

Hybrid
EUR 70,000 - 110,000
Hybrid remote work up to 3 days per
Offices with spacious, dynamic worksp​
Healthy meals served at headquarters
+2