Get more replies from employers
Send a job-specific resume in minutes.
Valtech is seeking a highly experienced Site Reliability Expert to lead observability, reliability, and operational excellence across complex cloud-native environments. The role combines production operations with automation and scalable monitoring in a distributed microservices setup.
The ideal candidate will drive SRE best practices, work with product teams, define observability standards, and enable robust monitoring across services. French and English communication is required.
We’retheexperience innovation company - a trusted partner to the world’s most recognized brands. To our people we offer growth opportunities, a values-driven culture, international careers and the chance to shape the future of experience. The opportunity At Valtech , you’ll find an environment designed for continuous learning, meaningful impact, and professional growth. Whether you're pioneering new digital solutions, challenging conventional thinking or building the next generation of customer experiences, your work will help transform industries.
Please be aware thet French speaking skills are needed for this role. We are seeking a highly experienced Site Reliability Expert to lead and drive observability, reliability, and operational excellence initiatives across complex cloud-native environments. This role goes beyond platform administration and requires a strong Site Reliability Engineering (SRE) background, combining observability expertise with production operations, automation, and reliability best practices. The ideal candidate will be an experienced technical leader capable of defining observability standards and strategies, supporting product teams, implementing reliability practices, and enabling scalable monitoring solutions across distributed microservices architectures.
You will thrive in this role if you are: A curious problem solver who challenges the status quo A collaborator who values teamwork and knowledge-sharing Excited by the intersection of technology, creativity and data Experienced in Agile methodologies and consulting (a plus)
To be considered for this role, you must meet the following essential qualifications: Site Reliability Engineering & Observability Significant experience in Site Reliability Engineering (SRE) within large-scale production environments. Deep understanding of: Service Level Indicators (SLIs) Service Level Objectives (SLOs) Error Budgets Symptom-based Alerting Proven expertise with enterprise observability platforms such as: Dynatrace Datadog New Relic AppDynamics Strong experience with: Application Performance Monitoring (APM) Real User Monitoring (RUM) Monitoring agents and instrumentation Alerting strategies Role-Based Access Control (RBAC) SLO management Tagging and governance models Distributed Systems & Cloud Platforms Strong knowledge of OpenTelemetry (OTEL) and distributed tracing. Experience working within composable, microservices-based architectures. Hands-on production experience with: AWS Kubernetes Automation & DevOps Experience with infrastructure and operational automation. Practical knowledge of: Terraform Bash scripting Python scripting Experience with CI/CD tools such as GitLab CI or equivalent pipeline/workflow platforms.
Demonstrated ability to lead technical initiatives and workstreams. Experience working within complex operational and agile environments. Strong stakeholder management and collaboration skills. Excellent communication skills in both French and English. Strong documentation practices, organizational skills, and attention to detail. High degree of autonomy and ownership.
Experience monitoring and supporting Java Spring Boot applications. Experience within e-commerce platforms and high-transaction environments. Experience establishing enterprise-wide observability frameworks and governance models. Consulting or advisory experience supporting multiple engineering teams.