Engineering Manager SRE - GPU Cloud

Scaleway

Paris

Hybride

EUR 110 000 - 160 000

Plein temps

14 jours+
Générateur de candidature

Une candidature complète en une minute — un CV et une lettre de motivation personnalisés, prêts à être envoyés.

Passez les filtres ATS

Avantages offerts par ce poste

Hybrid work 3 days remote
International environment
Healthy meals at HQ

Résumé du poste

Scaleway is seeking a Senior Site Reliability Engineer Lead to scale our GPU Cloud infrastructure. You will lead a team of 6 SREs, drive automation across bare-metal provisioning and lifecycle management, and own the technical roadmap for production GPU clusters.

You will work in a collaborative, international environment with hybrid work with up to 3 days remote per week, across offices in Paris and other French cities. Experience with Linux, Kubernetes and GPU/HPC infra is essential.

Qualifications

  • Strong experience managing SRE, Infrastructure or Platform Engineering teams in production environments.
  • Strong knowledge of Linux/UNIX systems and infrastructure fundamentals.
  • Good networking knowledge, including VLAN, VRF, VIP and NAT.
  • Experience with bare-metal infrastructure and remote server provisioning using technologies such as PXE, BMC or IPMI.
  • Strong understanding of infrastructure automation, observability and incident management.
  • Experience with Kubernetes and monitoring stacks such as Prometheus and Grafana.
  • Experience with GPU, HPC or high-performance infrastructure is a strong plus.

Responsabilités

  • Lead and manage a team of 6 Site Reliability Engineers, supporting both their technical execution and career development
  • Provide technical leadership and challenge architecture decisions related to large-scale GPU infrastructure
  • Design and drive automation for bare-metal server provisioning and lifecycle management
  • Improve remote deployment and server management capabilities using technologies and concepts such as PXE, BMC and IPMI
  • Drive automation around hardware failures, remediation and server recovery
  • Build and improve observability, monitoring, logging and alerting capabilities across production GPU clusters
  • Own the SRE team's technical roadmap, priorities and delivery
  • Ensure the reliability, scalability, performance and resilience of production GPU infrastructure
  • Drive continuous improvements in automation, incident response and post-incident remediation
  • Support technical decisions involving Linux systems, networking, hardware and cluster architecture
  • Collaborate closely with Engineering, Hardware, Product, Operations and other GPU Cloud teams
  • Recruit, onboard, coach and develop engineers within the team
  • Lead incident response and post-incident improvements when critical production issues occur

Connaissances

Linux/UNIX systems
VLAN/VRF/NAT
Bare-metal provisioning
PXE/BMC/IPMI
Kubernetes
Prometheus
Grafana
GPU/HPC infrastructure

Description du poste

OUR STORY:

Join Scaleway and shape the sovereign cloud of tomorrow !

Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies.

Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector.

With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants.

Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen.

Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.

WHY WE NEED YOU ?

As our GPU Cloud infrastructure continues to scale, we are strengthening our SRE organization to support the deployment and operation of increasingly large and complex AI and HPC infrastructure.

Your mission will be to lead our Site Reliability Engineering team and ensure the reliability, scalability, and operational excellence of our GPU clusters.

This is not a traditional IT Operations management role. You will combine engineering leadership with strong technical ownership, working close to the infrastructure itself from Linux systems, networking and bare-metal server provisioning to hardware lifecycle, automation and cluster observability.

You will help the team automate critical infrastructure workflows, improve reliability and operate production-grade GPU platforms powering our sovereign cloud.

YOUR FUTURE TEAM

We work in a collaborative and international environment where the diversity of Scalers, combined with a strong culture of knowledge sharing, helps us bring ambitious projects to life.

You will lead a team of 6 Site Reliability Engineers within the GPU Cloud organization.

The team works on some of our most critical AI and HPC infrastructure challenges, including bare-metal provisioning, GPU cluster automation, server lifecycle management, hardware failure management, observability, reliability and the integration of new GPU technologies.

The scope goes beyond traditional cloud-native infrastructure: the team operates close to the physical servers and needs to automate the full lifecycle of large fleets of GPU machines, from remote provisioning to production operations and remediation.

You will collaborate closely with GPU Cloud Engineering, Hardware, Product and Operations teams, as well as other infrastructure teams across Scaleway.

YOUR DAILY ROUTINE

Tasks

  • Lead and manage a team of 6 Site Reliability Engineers, supporting both their technical execution and career development
  • Provide technical leadership and challenge architecture decisions related to large-scale GPU infrastructure
  • Design and drive automation for bare-metal server provisioning and lifecycle management
  • Improve remote deployment and server management capabilities using technologies and concepts such as PXE, BMC and IPMI
  • Drive automation around hardware failures, remediation and server recovery
  • Build and improve observability, monitoring, logging and alerting capabilities across production GPU clusters
  • Own the SRE team's technical roadmap, priorities and delivery
  • Ensure the reliability, scalability, performance and resilience of production GPU infrastructure
  • Drive continuous improvements in automation, incident response and post-incident remediation
  • Support technical decisions involving Linux systems, networking, hardware and cluster architecture
  • Collaborate closely with Engineering, Hardware, Product, Operations and other GPU Cloud teams
  • Recruit, onboard, coach and develop engineers within the team
  • Lead incident response and post-incident improvements when critical production issues occur
ABOUT YOU

HARDSKILLS:

  • Strong experience managing SRE, Infrastructure or Platform Engineering teams in production environments
  • Strong knowledge of Linux/UNIX systems and infrastructure fundamentals
  • Good networking knowledge, including VLAN, VRF, VIP and NAT
  • Experience with bare-metal infrastructure and remote server provisioning using technologies such as PXE, BMC or IPMI
  • Strong understanding of infrastructure automation, observability and incident management
  • Experience with Kubernetes and monitoring stacks such as Prometheus and Grafana
  • Experience with GPU, HPC or high-performance infrastructure is a strong plus

SOFT SKILLS:

  • Strong engineering leadership and team management capabilities
  • Technical rigor and high attention to detail in production-critical environments
  • Ability to handle high-pressure operational situations and manage incident stress pragmatically
  • Excellent communication skills with the ability to convey challenging messages effectively
  • Collaborative mindset with a focus on empowering engineers rather than micromanaging
WHAT YOU WILL FIND AT SCALEWAY
  • Hybrid work: We offer up to 3 days of remote work per week.
  • Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
  • Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
  • Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.
  • International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.
  • Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.

Why join the Scaleway adventure?

A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.

A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.

Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.

THE NEXT STEPS …
  • Discovery call with HR
  • Technical interview with the HPC team to understand your technical skills and approach to the role
  • Manager interview to validate your expertise
  • Interview with an Engineering Manager / Head of Engineering to deepen discussions and assess your fit with the team
  • HR interview and office visit to tour our offices and meet your future colleagues

Scaleway is a company certified under SecNumCloud. A background check is mandatory in order to join the company.

At Scaleway, we are committed to building an inclusive and respectful workplace where everyone has a fair opportunity to thrive.

All applications are considered with care, regardless of age, gender, sexual orientation, ethnic or social background, religion, disability, or any other characteristic.

We believe great ideas come from everywhere, and everyone which is why you should definitely apply.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Head of Engineering - GPU Cloud
Head of Engineering - GPU Cloud

Scaleway • Paris

Hybride
EUR 140 000 - 190 000
Hybrid work
Office spaces
Meal service
+3
SRE Engineering Manager – GPU Cloud
SRE Engineering Manager – GPU Cloud

Webhosting • Paris

Hybride
EUR 90 000 - 130 000
Hybrid work
Dining service
Swile card
Head of Engineering - GPU Cloud
Head of Engineering - GPU Cloud

Scaleway • Paris

Sur place
EUR 180 000 - 240 000
Pre-Sales Solutions Architect - AI & GPU Infrastructure
Pre-Sales Solutions Architect - AI & GPU Infrastructure

Scaleway • Toulouse

Hybride
EUR 65 000 - 90 000
Hybrid work
Lunch service
Swile card for lunches
+2
Head of Operations GPU Cloud
Head of Operations GPU Cloud

Scaleway • Paris

Hybride
EUR 90 000 - 130 000
Hybrid work up to 3 days remote per wk
Modern offices near public transport
Healthy meal service at HQ
+3
Head of Operations GPU Cloud
Head of Operations GPU Cloud

Webhosting • Paris

Hybride
EUR 110 000 - 150 000
Hybrid work
Modern offices
Healthy meals
+3
Pre-Sales Solutions Architect - AI & GPU Infrastructure
Pre-Sales Solutions Architect - AI & GPU Infrastructure

Scaleway • Lille

Hybride
EUR 85 000 - 110 000
Hybrid work up to 3 days per week
International environment with diverse
Presales Solutions Engineer - HPC
Presales Solutions Engineer - HPC

Scaleway • France

Hybride
EUR 70 000 - 110 000
Hybrid work up to 3 days remote per wk
Office near public transport
Hardware Architect
Hardware Architect

Scaleway • Paris

Hybride
EUR 60 000 - 80 000
Healthy meal service
Remote work flexibility
Access to gym and daycare services
+1
Hardware Architect
Hardware Architect

Scaleway • Paris

Hybride
EUR 70 000 - 90 000
Hybrid work model
Healthy meal service
Access to a gym
+1