Senior Site Reliability Engineer (SRE, Compute Node Team)

Jobgether

France

Sur place

EUR 90 000 - 140 000

Plein temps

Il y a 3 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.

Passez les filtres ATS

Avantages offerts par ce poste

Competitive compensation
Learning opportunities
Significant ownership in your work
Collaborative engineering environment

Résumé du poste

Jobgether is seeking a Senior Site Reliability Engineer in France to own compute node reliability across large-scale virtualized workloads. You will work closely with Linux systems, virtualization stacks, and containerized workloads to improve performance, observability, and incident response.

The role emphasizes deep Linux expertise, QEMU/KVM experience, and collaboration with platform and infra teams to enhance compute platform reliability for AI and cloud workloads.

Qualifications

  • Significant professional experience in Site Reliability Engineering, Systems Engineering, Linux infrastructure, or a closely related field.
  • Deep expertise in Linux, including strong understanding of both user space and kernel space.
  • Hands-on experience with QEMU/KVM and virtualization technologies.
  • Experience operating production systems and responding to incidents.
  • Experience building and operating observability stacks and reliability signals.

Responsabilités

  • Ensure the reliability, availability, and performance of compute nodes running virtual machines.
  • Analyze and debug Linux systems across user space and kernel space.
  • Investigate system capabilities, dependencies, and trade-offs across the stack.
  • Troubleshoot production issues involving CPU, memory, NUMA, cgroups, and scheduling.
  • Work hands-on with virtualization technologies, primarily QEMU/KVM.
  • Design and evolve observability for the compute node layer with metrics, logs, traces, alerts, SLIs and SLOs.
  • Lead or contribute to incident response and postmortems for reliability improvements.
  • Collaborate with platform, kernel/hypervisor, GPU, and infra teams.

Connaissances

Linux
QEMU/KVM
Containers
Observability
Incident response
Root-cause analysis
SRE discipline
SLIs and SLOs
Kubernetes internals
Linux debugging tools

Outils

QEMU/KVM
perf
eBPF
ftrace
strace

Description du poste

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer (SRE, Compute Node Team) based in France.

This is a senior Site Reliability Engineering role focused on the infrastructure that runs and manages virtual machines across a large-scale cloud platform.
You will work close to the Linux operating system, hypervisor, and node-level services that form the foundation of the compute environment.
The role combines deep Linux systems engineering, virtualization, containerization, observability, and production reliability.
You will investigate complex issues involving CPU, memory, NUMA, cgroups, scheduling, and system performance across user and kernel space.
You will also help shape reliability practices through strong monitoring, incident response, root-cause analysis, and postmortem processes.
Collaboration with platform, kernel, hypervisor, GPU, and infrastructure teams will be central to improving system design and operability.
This is an opportunity to influence critical compute infrastructure supporting demanding AI and cloud workloads at significant scale.

Accountabilities
  • Ensure the reliability, availability, and performance of compute nodes responsible for running virtual machines.
  • Analyze and debug complex Linux systems across both user space and kernel space.
  • Investigate system capabilities, limitations, dependencies, and trade-offs across different layers of the operating system and infrastructure stack.
  • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups, and scheduling.
  • Work hands-on with virtualization technologies, primarily QEMU/KVM and Linux-native technologies.
  • Analyze VM lifecycle behavior, performance characteristics, resource utilization, and failure modes.
  • Support and improve containerized workloads using Linux-native mechanisms such as namespaces and cgroups.
  • Design and evolve observability for the compute node layer, including metrics, logs, traces, alerts, SLIs, and SLOs.
  • Build reliability signals that provide clear and actionable insight into system behavior.
  • Lead or contribute to incident response, ensuring production issues are diagnosed and resolved efficiently.
  • Conduct structured root-cause analysis and develop corrective actions for recurring or systemic reliability issues.
  • Lead and contribute to postmortems focused on long-term reliability improvements rather than short-term remediation alone.
  • Identify opportunities to automate operational processes and improve the resilience of compute infrastructure.
  • Collaborate closely with platform, kernel/hypervisor, GPU, and infrastructure teams on system design and operational improvements.
  • Contribute to improving the operability, scalability, and maintainability of node-level services.
  • Investigate performance issues across multiple layers of the compute stack and develop practical engineering solutions.
  • Help establish reliability and observability as core capabilities of the compute platform.
Requirements:
  • Significant professional experience in Site Reliability Engineering, Systems Engineering, Linux infrastructure, or a closely related field.
  • Deep expertise in Linux, including strong understanding of both user space and kernel space.
  • Knowledge of important Linux kernel subsystems, including scheduling, memory management, filesystems, cgroups, and namespaces.
  • Strong understanding of system boundaries, constraints, dependencies, and trade-offs across different infrastructure layers.
  • Hands‑on experience with QEMU/KVM and a solid understanding of virtualization technologies.
  • Understanding of virtual machine lifecycles, performance characteristics, resource management, and failure modes.
  • Practical experience with containers, Linux namespaces, and cgroups.
  • Strong understanding of resource isolation, allocation, and control in containerized environments.
  • Excellent debugging skills and the ability to reason systematically about complex system failures.
  • Structured, hypothesis‑driven approach to incident investigation and troubleshooting.
  • Strong understanding of the SRE discipline, including the relationship between software engineering, operations, reliability, and system design.
  • Experience building and operating observability stacks, rather than simply consuming existing monitoring dashboards.
  • Ability to translate complex system behavior into actionable reliability signals, alerts, SLIs, and SLOs.
  • Strong analytical and problem‑solving skills, with the ability to investigate issues across operating‑system and infrastructure layers.
  • Experience operating production systems and responding effectively to reliability and performance incidents.
  • Strong communication and collaboration skills when working with multidisciplinary infrastructure and engineering teams.
  • Ability to take ownership of complex technical problems and drive them through investigation, resolution, and long‑term improvement.
  • Experience with Kubernetes internals or node‑level components is an advantage.
  • Hands‑on experience with low‑level Linux debugging tools such as perf, eBPF, ftrace, strace, or kernel crash dumps is beneficial.
  • Familiarity with large‑scale compute or bare‑metal infrastructure is a plus.
  • Contributions to open‑source infrastructure or systems software are advantageous.
  • Experience debugging hardware‑and driver‑level issues, including GPUs, NVLink, or InfiniBand, is a strong plus.
Benefits:
  • Competitive compensation.
  • Career growth and continuous learning opportunities.
  • Flexibility and significant ownership in your work.
  • Collaborative and innovative engineering environment.
  • Opportunity to work on impactful AI and cloud infrastructure projects.
  • Exposure to large-scale compute, Linux systems, virtualization, and distributed infrastructure.
  • Opportunity to collaborate with highly skilled international engineering teams.
  • Inclusive workplace committed to equal employment opportunities.
  • Workplace accommodations available during the application process where required.
  • Employment is subject to authorization to work in the country where the position is based.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre‑contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior SRE - Compute Node Linux Reliability
Senior SRE - Compute Node Linux Reliability

Jobgether • France

Sur place
EUR 90 000 - 140 000
Competitive compensation
Learning opportunities
Significant ownership in your work
+1
Site Reliability Engineer (SRE) - AI GPU Clusters
Site Reliability Engineer (SRE) - AI GPU Clusters

Scaleway • Paris

Sur place
EUR 50 000 - 70 000
Hybrid work: up to 3 remote days per week
Chef-served meals
Access to gym and daycare
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • France

Sur place
EUR 110 000 - 135 000
Competitive compensation
International teams
Ownership of projects
+2
Engineering Manager SRE - GPU Cloud
Engineering Manager SRE - GPU Cloud

Scaleway • Paris

Hybride
EUR 110 000 - 160 000
Hybrid work 3 days remote
International environment
Healthy meals at HQ
SRE Engineering Manager – GPU Cloud
SRE Engineering Manager – GPU Cloud

Webhosting • Paris

Hybride
EUR 90 000 - 130 000
Hybrid work
Dining service
Swile card
Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral • Paris

Sur place
EUR 90 000 - 130 000
Site Reliability Engineer - Network
Site Reliability Engineer - Network

Scaleway • Paris

Hybride
EUR 40 000 - 70 000
Hybrid work: Up to 3 days remote
Healthy meal service
Spacious, dynamic workspaces
+2
Site Reliability Engineer - SRE Paris
Site Reliability Engineer - SRE Paris

Scaleway • Toulouse

Sur place
EUR 80 000 - 110 000
Site Reliability Engineer - Network
Site Reliability Engineer - Network

Scaleway • Paris

Hybride
EUR 45 000 - 70 000
Télétravail jusqu'à 3 jours par semaine
Espaces de travail spacieux et dynamiques
Service de repas sain au siège
+3
Site Reliability Engineer
Site Reliability Engineer

Helsing • Paris

Hybride
EUR 90 000 - 130 000
Competitive compensation
VSOP options
Relocation support
+3