Une candidature complète en une minute — un CV et une lettre de motivation personnalisés, prêts à être envoyés.
Scaleway is seeking a Senior Site Reliability Engineer Lead to scale our GPU Cloud infrastructure. You will lead a team of 6 SREs, drive automation across bare-metal provisioning and lifecycle management, and own the technical roadmap for production GPU clusters.
You will work in a collaborative, international environment with hybrid work with up to 3 days remote per week, across offices in Paris and other French cities. Experience with Linux, Kubernetes and GPU/HPC infra is essential.
OUR STORY:
Join Scaleway and shape the sovereign cloud of tomorrow !
Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies.
Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector.
With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants.
Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen.
Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.
As our GPU Cloud infrastructure continues to scale, we are strengthening our SRE organization to support the deployment and operation of increasingly large and complex AI and HPC infrastructure.
Your mission will be to lead our Site Reliability Engineering team and ensure the reliability, scalability, and operational excellence of our GPU clusters.
This is not a traditional IT Operations management role. You will combine engineering leadership with strong technical ownership, working close to the infrastructure itself from Linux systems, networking and bare-metal server provisioning to hardware lifecycle, automation and cluster observability.
You will help the team automate critical infrastructure workflows, improve reliability and operate production-grade GPU platforms powering our sovereign cloud.
We work in a collaborative and international environment where the diversity of Scalers, combined with a strong culture of knowledge sharing, helps us bring ambitious projects to life.
You will lead a team of 6 Site Reliability Engineers within the GPU Cloud organization.
The team works on some of our most critical AI and HPC infrastructure challenges, including bare-metal provisioning, GPU cluster automation, server lifecycle management, hardware failure management, observability, reliability and the integration of new GPU technologies.
The scope goes beyond traditional cloud-native infrastructure: the team operates close to the physical servers and needs to automate the full lifecycle of large fleets of GPU machines, from remote provisioning to production operations and remediation.
You will collaborate closely with GPU Cloud Engineering, Hardware, Product and Operations teams, as well as other infrastructure teams across Scaleway.
Tasks
HARDSKILLS:
SOFT SKILLS:
Why join the Scaleway adventure?
A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.
A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.
Scaleway is a company certified under SecNumCloud. A background check is mandatory in order to join the company.
At Scaleway, we are committed to building an inclusive and respectful workplace where everyone has a fair opportunity to thrive.
All applications are considered with care, regardless of age, gender, sexual orientation, ethnic or social background, religion, disability, or any other characteristic.
We believe great ideas come from everywhere, and everyone which is why you should definitely apply.