Staff Site Reliability Engineer

Wand AI

Brussel

Sur place

EUR 80 000 - 110 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

Wand AI is seeking a Senior Staff SRE Engineer located in Brussels, Belgium, to act as a senior technical authority within our reliability function. This deeply hands-on role involves designing resilient infrastructure, ensuring our AI-driven products maintain high availability and performance, and collaborating with product and engineering teams.

The ideal candidate will have extensive experience in cloud infrastructure, particularly with AWS or Azure, and proficiency in Kubernetes and CI/CD practices. Mentoring engineers and enhancing reliability standards is key to this position.

Qualifications

  • Extensive hands-on experience in Site Reliability Engineering or Production Engineering roles.
  • Strong experience with Kubernetes including migration and scaling.
  • Experience supporting production data platforms and ML systems.

Responsabilités

  • Architect, deploy, and operate scalable production environments.
  • Lead reliability improvements across engineering streams.
  • Mentor engineers to enhance overall reliability across teams.

Connaissances

Cloud Infrastructure Expertise
Kubernetes Management
CI/CD Pipeline Optimization
Infrastructure as Code
Observability Tools Experience

Outils

Terraform
AWS
Azure

Description du poste

Position Summary

We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands‑on individual contributor role, to build and operate SRE practices at scale. You will design and evolve resilient infrastructure, drive reliability across multiple engineering streams, and ensure our AI‑driven products operate with high availability, performance, and security. You will work across platform, product, data, and ML teams, helping us productionise models, absorb and standardise customer environments, strengthen Kubernetes‑based architecture, and mature our CI/CD pipelines end‑to‑end. You will also collaborate with other Staff engineers and Architects to shape the global product architect and technology vision.

Responsibilities
  • Architect, deploy, and operate scalable, secure production environments (AWS preferred
  • Lead reliability improvements across multiple engineering streams.
  • Design and evolve Kubernetes‑based infrastructure, including migration and optimisation initiatives.
  • Build and enforce strong Infrastructure‑as‑Code standards.
  • Define and operationalise SLIs, SLOs, and error budgets.
  • Strengthen observability across applications, infrastructure, data pipelines, and ML systems.
  • Work closely with product and data teams to integrate model analytics and product telemetry into reliability insights.
  • Work across and optimise the entire CI/CD pipeline, from build to deploy to rollback.
  • Improve release safety, deployment frequency, and predictability of SLAs.
  • Lead incident response for complex cross‑system failures and drive postmortems.
  • Reduce operational toil through automation and platform engineering improvements.
  • Design processes and tooling to absorb, standardise, and troubleshoot customer environments.
  • Support and productionise ML workloads (MLOps practices including model deployment, monitoring, retraining workflows).
  • Ensure infrastructure aligns with enterprise‑grade security and regulatory requirements.
  • Mentor engineers and raise the overall reliability bar across teams.
Key Requirements
  • Extensive hands‑on experience in SRE or Production Engineering roles.
  • Demonstrated experience building or scaling SRE practices in high‑growth or complex environments.
  • Deep expertise in AWS or Azure‑based cloud infrastructure.
  • Strong experience with Kubernetes (including migration, scaling, and production hardeni
  • ng). Advanced Infrastructure‑as‑Code experience (Terraform or equivalent).
  • End‑to‑end CI/CD pipeline design and optimisation experience.
  • Strong experience with observability tooling across distributed systems.
  • Experience troubleshooting complex multi‑tenant or customer‑hosted environments.
  • Experience supporting production data platforms and ML systems.
  • MLOps experience, including model deployment and monitoring.
  • Strong understanding of distributed systems, scalability, and fault tolerance.
  • Systems thinker who understands interactions across infrastructure, product, data, and ML.
  • Excellent communication skills and ability to work cross‑functionally.
Preferred Experience
  • Experience in large‑scale global B2B/B2C products.
  • Experience working with AI/ML systems, NLP, or LLM‑based products.
  • Experience integrating product analytics and model performance metrics into operational monitoring.
  • Background in enterprise environments with strong security and compliance requirements.
  • Experience implementing regulatory controls within cloud infrastructure.
  • Experience evaluating infrastructure tooling and vendors.
  • Experience in collaborating with large scale enterprise customers to deploy and operate environments within their accounts and VPCs.
Personal Characteristics
  • Strong problem solver who anticipates failure modes.
  • High ownership mentality and accountability.
  • Comfortable working across streams and influencing without formal authority.
  • Learning‑oriented with a drive for continuous improvement.
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior Staff SRE: Resilient Cloud, Kubernetes & MLOps
Senior Staff SRE: Resilient Cloud, Kubernetes & MLOps

Wand AI • Brussel

Sur place
EUR 80 000 - 110 000
Head of Infrastructure & Security
Head of Infrastructure & Security

Deliverect • Gent

Sur place
EUR 90 000 - 120 000
SRE Engineer | fiabilité de plateformes à l'internationale
SRE Engineer | fiabilité de plateformes à l'internationale

Editx • Oudergem

Sur place
EUR 50 000 - 70 000
Site Reliability Engineering (SRQ159374)
Site Reliability Engineering (SRQ159374)

TiTANS Consulting • Brussel Hoofdstad

Sur place
EUR 55 000 - 75 000
Senior SRE: Platform Reliability & Observability (Remote EU)
Senior SRE: Platform Reliability & Observability (Remote EU)

Koda Tech • Belgique

Sur place
EUR 70 000 - 110 000
Site Reliability Engineer
Site Reliability Engineer

Davinsi Labs • Berchem

Sur place
EUR 75 000 - 110 000
Bonus
Medical coverage
Flexible mobility options
+3
Site Reliability Engineer
Site Reliability Engineer

Qargo TMS • Gent

Hybride
EUR 80 000 - 120 000
Flexible hours
Hybrid working
Green office
+2
Reliability Engineer
Reliability Engineer

CBRE Group, Inc. • Oost-Vlaanderen

Sur place
EUR 45 000 - 60 000
Cloud Engineer (SRE – Dynatrace, Cloud Observability & Security)
Cloud Engineer (SRE – Dynatrace, Cloud Observability & Security)

afarax • Brussel Hoofdstad

Sur place
EUR 111 000 - 166 000
Extensive project network
Access to relevant client projects
Support throughout the engagement
+1
Site Reliability Engineer
Site Reliability Engineer

JobDev • Brussel Hoofdstad

Hybride
EUR 70 000 - 90 000
Hybrid work environment
Learning & growth opportunities
Ownership in infrastructure design