Site Reliability Expert

NEPSE Trading

Canada

On-site

CAD 100,000 - 150,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive insurance plan (Gold/SIl
Virtual healthcare via Sun Life
Employee and Family Assistance Program
Mental health support program
Personal Spending Account $500

Job summary

Valtech is seeking a highly experienced Site Reliability Expert to lead observability, reliability, and operational excellence across complex cloud-native environments. The role combines production operations with automation and scalable monitoring in a distributed microservices setup.

The ideal candidate will drive SRE best practices, work with product teams, define observability standards, and enable robust monitoring across services. French and English communication is required.

Qualifications

  • Significant experience in Site Reliability Engineering (SRE) within large-scale production environments.
  • Deep understanding of SLIs, SLOs, and error budgets; robust symptom-based alerting practices.
  • Strong RBAC and governance modelling; tagging and ownership standards.
  • Experience with enterprise observability platforms and distributed tracing in microservices.

Responsibilities

  • Define and implement observability strategies, governance, and standards across applications and platforms.
  • Design and maintain monitoring, alerting, dashboards, and reporting using Dynatrace or equivalent tools.
  • Establish SRE best practices and drive reliability metrics across teams and products.
  • Collaborate with engineering and product teams to improve system reliability and performance.
  • Lead technical workstreams, prioritize initiatives, and ensure on-time delivery within budgets.

Skills

SRE & Observability
SLI/SLO/Error budgets
RBAC & governance
Distributed systems & microservices
Cloud platforms (AWS)
Kubernetes
Automation & DevOps
Agile methodologies

Tools

Dynatrace
Datadog
New Relic
AppDynamics
OpenTelemetry

Job description

Why Valtech ?

We’retheexperience innovation company - a trusted partner to the world’s most recognized brands. To our people we offer growth opportunities, a values-driven culture, international careers and the chance to shape the future of experience. The opportunity At Valtech , you’ll find an environment designed for continuous learning, meaningful impact, and professional growth. Whether you're pioneering new digital solutions, challenging conventional thinking or building the next generation of customer experiences, your work will help transform industries.

  • The work we do and the innovation we drive
  • Our values of share, care and dare
  • A workplace culture that fosters creativity, diversity and autonomy
  • Our borderless, global framework, which enables seamless collaboration
The role

Please be aware thet French speaking skills are needed for this role. We are seeking a highly experienced Site Reliability Expert to lead and drive observability, reliability, and operational excellence initiatives across complex cloud-native environments. This role goes beyond platform administration and requires a strong Site Reliability Engineering (SRE) background, combining observability expertise with production operations, automation, and reliability best practices. The ideal candidate will be an experienced technical leader capable of defining observability standards and strategies, supporting product teams, implementing reliability practices, and enabling scalable monitoring solutions across distributed microservices architectures.

You will thrive in this role if you are: A curious problem solver who challenges the status quo A collaborator who values teamwork and knowledge-sharing Excited by the intersection of technology, creativity and data Experienced in Agile methodologies and consulting (a plus)

Role responsibilities
  • Define and implement observability strategies, standards, and governance across applications and platforms.
  • Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms.
  • Establish and drive SRE best practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting.
  • Partner with engineering and product teams to improve system reliability, performance, and operational maturity.
  • Develop standards for tagging, ownership, dashboard design, access management, and alerting governance.
  • Support teams that are not specialized in observability by providing guidance, coaching, and knowledge transfer.
  • Lead technical workstreams, prioritize initiatives, and ensure successful delivery within defined timelines and budgets.
  • Analyze distributed systems and troubleshoot complex production issues using monitoring and tracing data.
  • Promote documentation, operational rigor, and continuous improvement across engineering teams.
  • Collaborate effectively within a distributed, multilingual environment.
Must have qualifications

To be considered for this role, you must meet the following essential qualifications: Site Reliability Engineering & Observability Significant experience in Site Reliability Engineering (SRE) within large-scale production environments. Deep understanding of: Service Level Indicators (SLIs) Service Level Objectives (SLOs) Error Budgets Symptom-based Alerting Proven expertise with enterprise observability platforms such as: Dynatrace Datadog New Relic AppDynamics Strong experience with: Application Performance Monitoring (APM) Real User Monitoring (RUM) Monitoring agents and instrumentation Alerting strategies Role-Based Access Control (RBAC) SLO management Tagging and governance models Distributed Systems & Cloud Platforms Strong knowledge of OpenTelemetry (OTEL) and distributed tracing. Experience working within composable, microservices-based architectures. Hands-on production experience with: AWS Kubernetes Automation & DevOps Experience with infrastructure and operational automation. Practical knowledge of: Terraform Bash scripting Python scripting Experience with CI/CD tools such as GitLab CI or equivalent pipeline/workflow platforms.

Delivery & Collaboration

Demonstrated ability to lead technical initiatives and workstreams. Experience working within complex operational and agile environments. Strong stakeholder management and collaboration skills. Excellent communication skills in both French and English. Strong documentation practices, organizational skills, and attention to detail. High degree of autonomy and ownership.

Nice to have qualifications

Experience monitoring and supporting Java Spring Boot applications. Experience within e-commerce platforms and high-transaction environments. Experience establishing enterprise-wide observability frameworks and governance models. Consulting or advisory experience supporting multiple engineering teams.

The benefits
  • This is a Full time position based in Canada .
  • The offered salary range is $100,000 - $150,000 CAD annually, depending on experience and location.
  • A comprehensive insurance plan , where you can choose the module that best suits your needs—Gold, Silver, or Bronze.
  • The employer may contribute up to 80% of your coverage depending on the selected module.
  • This plan includes short- and long-term disability coverage .
  • Dialogue via Sun Life provides virtual healthcare services, allowing you to consult with a healthcare professional for emergencies, prescription renewals, and more.
  • You also have access to the Employee and Family Assistance Program , as well as a complete mental health support program .
  • A $500 Personal Spending Account , which can be used
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Expert
Site Reliability Expert

Valtech • Canada

On-site
CAD 100,000 - 150,000
Insurance plan
Retirement plan
Flexible vacation
+1
Site Reliability Expert
Site Reliability Expert

Valtech • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Assurance complète
REER et RPDB
Vacances flexibles
+2
Site Reliability Engineer
Site Reliability Engineer

Tecsys Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 120,000
Digital-first work environment
Collaborative workspaces
Continuous learning opportunities
Senior Site Reliability Engineer (SRE) – Automation & Observability
Senior Site Reliability Engineer (SRE) – Automation & Observability

Tech Talent International • Montreal (administrative region)

Hybrid
CAD 110,000 - 120,000
9% bonus
3–5 weeks vacation
RRSP contribution
+2
SRE specialist
SRE specialist

Intact Financial Corporation • Montreal (administrative region)

On-site
CAD 109,000 - 135,000
Flexible work arrangements
Possibility to purchase up to 5 extra days off
Wellness benefits including telemedicine
+1
DevOps Engineer/ Dynatrace
DevOps Engineer/ Dynatrace

Motion Recruitment • Toronto

On-site
CAD 110,000 - 140,000
Medical, Dental, and Vision Insurance
Vacation Time
Observability / DevOps Advisor
Observability / DevOps Advisor

Intact Financial Corporation • Montreal (administrative region)

Hybrid
CAD 101,000 - 125,000
Flexible work arrangements
Hybrid work model
Health and wellness benefits
+1
Solutions Engineer (Remote - Montreal)
Solutions Engineer (Remote - Montreal)

Dynatrace LLC • Montreal (administrative region)

Hybrid
CAD 110,000 - 160,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Akkodis • Toronto

Hybrid
CAD 140,000 - 165,000
Bonus
Benefits
Manager, Site Reliability Engineering (SRE)
Manager, Site Reliability Engineering (SRE)

Quantum Technology Recruiting Inc. (QTR) • Toronto

On-site
CAD 155,000 - 165,000