Senior SRE - Compute Node Linux Reliability

Jobgether

France

Sur place

EUR 90 000 - 140 000

Plein temps

Il y a 3 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Démarquez-vous pour ce poste — générez un CV et une lettre de motivation personnalisés en environ une minute.

Passez les filtres ATS

Avantages offerts par ce poste

Competitive compensation
Learning opportunities
Significant ownership in your work
Collaborative engineering environment

Résumé du poste

Jobgether is seeking a Senior Site Reliability Engineer in France to own compute node reliability across large-scale virtualized workloads. You will work closely with Linux systems, virtualization stacks, and containerized workloads to improve performance, observability, and incident response.

The role emphasizes deep Linux expertise, QEMU/KVM experience, and collaboration with platform and infra teams to enhance compute platform reliability for AI and cloud workloads.

Qualifications

  • Significant professional experience in Site Reliability Engineering, Systems Engineering, Linux infrastructure, or a closely related field.
  • Deep expertise in Linux, including strong understanding of both user space and kernel space.
  • Hands-on experience with QEMU/KVM and virtualization technologies.
  • Experience operating production systems and responding to incidents.
  • Experience building and operating observability stacks and reliability signals.

Responsabilités

  • Ensure the reliability, availability, and performance of compute nodes running virtual machines.
  • Analyze and debug Linux systems across user space and kernel space.
  • Investigate system capabilities, dependencies, and trade-offs across the stack.
  • Troubleshoot production issues involving CPU, memory, NUMA, cgroups, and scheduling.
  • Work hands-on with virtualization technologies, primarily QEMU/KVM.
  • Design and evolve observability for the compute node layer with metrics, logs, traces, alerts, SLIs and SLOs.
  • Lead or contribute to incident response and postmortems for reliability improvements.
  • Collaborate with platform, kernel/hypervisor, GPU, and infra teams.

Connaissances

Linux
QEMU/KVM
Containers
Observability
Incident response
Root-cause analysis
SRE discipline
SLIs and SLOs
Kubernetes internals
Linux debugging tools

Outils

QEMU/KVM
perf
eBPF
ftrace
strace

Description du poste

Jobgether is seeking a Senior Site Reliability Engineer in France to own compute node reliability across large-scale virtualized workloads. You will work closely with Linux systems, virtualization stacks, and containerized workloads to improve performance, observability, and incident response.

The role emphasizes deep Linux expertise, QEMU/KVM experience, and collaboration with platform and infra teams to enhance compute platform reliability for AI and cloud workloads.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.