SRE Engineering Manager – GPU Cloud

Webhosting

Paris

On-site

EUR 90,000 - 130,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid work
Dining service
Swile card

Job summary

Scaleway in Paris is seeking a senior Site Reliability Engineer to lead the GPU Cloud SRE team and build a production-grade infra powering our sovereign cloud. You will scale automated solutions across GPU clusters and ensure high availability.

You will collaborate with software engineering and product teams, mentor engineers, and contribute to the reliability roadmap. The role offers hybrid work with flexible remote days and offices near public transport.

Qualifications

  • Experience leading a 6-person SRE team.
  • Strong background in production-grade GPU infrastructure.
  • Proficiency with observability, logging and monitoring at scale.
  • Ability to design and implement automated server lifecycle solutions.

Responsibilities

  • Lead and manage a team of 6 Site Reliability Engineers.
  • Design automated solutions for server lifecycle management across GPU clusters.
  • Design and implement observability, logging, and monitoring solutions for large-scale GPU clusters.
  • Plan, prioritize, and manage the technical development roadmap for the SRE team.
  • Collaborate with software engineering, product, and cross-functional teams across Scaleway.
  • Handle recruitment and career management for team members.
  • Maintain, scale, and optimize high-availability production systems under heavy load.
  • Participate in on-call rotations to ensure production reliability and fast incident resolution.

Skills

Team leadership
SRE
GPU
Automation

Tools

Kubernetes
Monitoring stacks
Automation tools

Job description

Job Description
OUR STORY:

Join Scaleway and shape the sovereign cloud of tomorrow ! Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies. Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector. With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants. Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen. Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.

WHY WE NEED YOU ?

Our growth is driving us to strengthen our GPU Cloud team to support our expanding infrastructure and key AI roadmap initiatives.

Your mission will be leading the Site Reliability Engineering (SRE) team in order to build, automate, and maintain a highly reliable, production-grade GPU cluster infrastructure powering our sovereign cloud.

YOUR FUTURE TEAM

We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together.

You will be part of a team of 6 SREs within the GPU Cloud organization. The team focuses on critical AI and HPC infrastructure challenges, including automating key components of our stack and implementing support for modern GPU technologies.

YOUR DAILY ROUTINE
Tasks
  • Lead and manage a team of 6 Site Reliability Engineers, supporting their career growth and technical execution
  • Design and implement automated solutions for server lifecycle management across GPU clusters
  • Design and implement observability, logging, and monitoring solutions for large-scale GPU clusters
  • Plan, prioritize, and manage the technical development roadmap for the SRE team
  • Collaborate and coordinate closely with software engineering, product, and cross-functional teams across Scaleway
  • Handle recruitment and career management for team members
  • Maintain, scale, and optimize high-availability production systems under heavy load
  • Participate in on-call rotations to ensure production reliability and fast incident resolution
HARDSKILLS:
SOFT SKILLS:
WHAT YOU WILL FIND AT SCALEWAY ++++

Why join the Scaleway adventure?

Benefits
  • A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.
  • A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
  • Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.
  • Hybrid work:We offer up to 3 days of remote work per week.
  • Offices:Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
  • Dining:Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
  • Well-being commitments:Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.
  • International environment:With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.
  • Career & Mobility:Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.
  • Discovery call with HR
  • Technical interview with the HPC team to understand your technical skills and approach to the role
  • Manager interview to validate your expertise
  • Interview with an Engineering Manager / Head of Engineering to deepen discussions and assess your fit with the team
  • HR interview and office visit to tour our offices and meet your future colleagues
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Head of Engineering - GPU Cloud
Head of Engineering - GPU Cloud

Scaleway • Paris

On-site
EUR 180,000 - 240,000
Head of Operations GPU Cloud
Head of Operations GPU Cloud

Webhosting • Paris

On-site
EUR 110,000 - 150,000
Hybrid work
Modern offices
Healthy meals
+3
Site Reliability Engineer (SRE) - AI GPU Clusters
Site Reliability Engineer (SRE) - AI GPU Clusters

Scaleway • Paris

On-site
EUR 50,000 - 70,000
Hybrid work: up to 3 remote days per week
Chef-served meals
Access to gym and daycare
Presales Solutions Engineer - HPC
Presales Solutions Engineer - HPC

Scaleway • France

On-site
EUR 70,000 - 110,000
Hybrid work up to 3 days remote per wk
Office near public transport
Hardware Enablement Engineer (Linux)
Hardware Enablement Engineer (Linux)

Scaleway • Paris

On-site
EUR 45,000 - 70,000
Hybrid work model
Healthy meal service
Access to gym and daycare services
+1
Full Stack Software Engineer (Python / React)
Full Stack Software Engineer (Python / React)

Scaleway • Paris

On-site
EUR 50,000 - 70,000
Up to 3 days remote work per week
Healthy meal service
Access to gym
+1
Site Reliability Engineer - SRE
Site Reliability Engineer - SRE

Scaleway • Lyon

Hybrid
EUR 75,000 - 110,000
Hybrid work (up to 3 days remote)
Lunch card (Swile)
Gym access
+1
Site Reliability Engineer - SRE
Site Reliability Engineer - SRE

Scaleway • Bordeaux

Hybrid
EUR 70,000 - 110,000
Hybrid remote work up to 3 days per
Offices with spacious, dynamic worksp​
Healthy meals served at headquarters
+2
Presales Solutions Engineer – HPC
Presales Solutions Engineer – HPC

Webhosting • Toulouse

On-site
EUR 90,000 - 130,000
Hybrid work (up to 3 days remote per w
Modern offices near public transport
Developer and tech community events
DevOps Cybersecurity Engineer
DevOps Cybersecurity Engineer

Scaleway • Rouen

On-site
EUR 70,000 - 110,000
Hybrid work up to 3 days per week
Modern offices near public transport
Healthy meals at HQ
+1