Site Reliability Engineer (SRE) - AI GPU Clusters

Scaleway

Paris

Sur place

EUR 50 000 - 70 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Hybrid work: up to 3 remote days per week
Chef-served meals
Access to gym and daycare

Résumé du poste

Scaleway in Paris is looking for a Site Reliability Engineer to build and support a vast AI infrastructure. The successful candidate will collaborate with engineering teams to tackle production challenges and ensure service continuity.

A hybrid work model allows up to three remote workdays per week. Scaleway offers various benefits, including healthy meals, gym access, and opportunities for career mobility within Iliad Group.

Qualifications

  • Experience with Python, Go, or C++.
  • Hands-on experience with Linux systems (Ubuntu/Debian).
  • Familiarity with monitoring and logging tools.

Responsabilités

  • Build AI infrastructure with monitoring and remediation.
  • Troubleshoot production issues with engineering teams.
  • Participate in on-call rotation for incident handling.

Connaissances

Automation passion
Proactive mindset
Strong collaboration skills
Python proficiency
Scripting skills
Linux experience

Outils

Prometheus
Grafana
Ansible
MariaDB

Description du poste

OUR STORY

Since 1999, Scaleway has designed secure and sustainable infrastructures for ambitious companies. In 2015, we shifted to cloud computing, becoming a leading European cloud provider. With AI investments, we are building sovereign AI alternatives. Our products serve a diverse range of customers across France and globally.

WHY WE NEED YOU?

Our growth drives us to strengthen our SRE team to support and scale production environments.

YOUR DAILY ROUTINE
  • Build a large AI infrastructure with monitoring, diagnosis, and remediation.
  • Troubleshoot high-impact production issues with engineering teams.
  • Participate in on‑call rotation to handle incidents and ensure service continuity.
  • Implement and maintain observability solutions for AI infrastructure and application health.
  • Contribute to AI infrastructure lifecycle management across environments and countries.
  • Promote and apply best practices in stability, resiliency, scalability, and security.
  • Maintain clear technical documentation for tools and procedures.
  • Collaborate closely with development teams to ensure infrastructure readiness.
  • Participate in team rituals and knowledge‑sharing initiatives.
ABOUT YOU
SOFT SKILLS
  • Proactive and solution-oriented mindset.
  • Passion for automation and continuous improvement.
  • Strong collaboration and communication skills.
  • Ability to work independently and in a team.
  • Willingness to mentor and share knowledge.
HARD SKILLS
  • Experience with Python, Go, or C++.
  • Strong scripting skills (Bash, Python).
  • Hands-on experience with Linux systems (Ubuntu/Debian).
  • Preferred hands-on experience with GPU & HPC infrastructure.
  • Knowledge of networking (TCP/IP, DNS, BGP, load-balancing, IPv6, etc.).
  • Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.).
  • Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.).
  • Experience managing relational databases (MariaDB).
  • Understanding of CI/CD pipelines (GitLab).
  • Comfortable with English (written and spoken).
BENEFITS AT SCALEWAY

Hybrid work: up to 3 days remote per week.

Offices: spacious, dynamic with outdoor spaces and bike parking.

Dining: chef-served healthy meals at headquarters; breakfast available all sites; Swile card for lunches at regional sites.

Well-being commitments: access to gym, daycare places, discounted services.

International environment: English widely spoken, diverse nationalities.

Career & Mobility: managers value internal mobility; opportunities to transition within Iliad Group.

INCLUSION STATEMENT

At Scaleway, we are committed to building an inclusive and respectful workplace where everyone has a fair opportunity to thrive. All applications are considered with care, regardless of age, gender, sexual orientation, ethnicity, religion, disability, or any other characteristic.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

SRE Engineering Manager – GPU Cloud
SRE Engineering Manager – GPU Cloud

Webhosting • Paris

Hybride
EUR 90 000 - 130 000
Hybrid work
Dining service
Swile card
Engineering Manager AI GPU Cloud
Engineering Manager AI GPU Cloud

Scaleway • Paris

Hybride
EUR 90 000 - 130 000
Hybrid work up to 3 days per week
Modern offices near public transport
Healthy meals at HQ and Swile lunchカード
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Toulouse

Hybride
EUR 70 000 - 110 000
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Lille

Hybride
EUR 70 000 - 90 000
Up to 3 days remote per week
Office near public transport
Swile meal card
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Rouen

Hybride
EUR 70 000 - 100 000
Hybrid work
Lunch service
Swile lunch card
+2
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Rennes

Hybride
EUR 65 000 - 90 000
Hybrid work
Free meals
Swile card
+3
Site Reliability Engineer (SRE) - Network Products
Site Reliability Engineer (SRE) - Network Products

Scaleway • Lyon

Hybride
EUR 70 000 - 100 000
Hybrid work (up to 3 days remote)
Office near public transport
Healthy meals at HQ
+4
Engineering Manager AI GPU Cloud
Engineering Manager AI GPU Cloud

Scaleway • Paris

Hybride
EUR 110 000 - 140 000
Hybrid work up to 3 days remote per wk
Office near public transport
Swile meal card
Pre-Sales Solutions Architect - AI & GPU Infrastructure
Pre-Sales Solutions Architect - AI & GPU Infrastructure

Scaleway • Lille

Hybride
EUR 85 000 - 110 000
Hybrid work up to 3 days per week
International environment with diverse
Head of Engineering - GPU Cloud
Head of Engineering - GPU Cloud

Scaleway • Paris

Hybride
EUR 140 000 - 190 000
Hybrid work
Office spaces
Meal service
+3