Senior Site Reliability Developer

Vena

Toronto

On-site

CAD 123,000 - 167,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Vena is seeking a Senior Site Reliability Developer to design and operate highly scalable, fault-tolerant services across AWS and Azure. You will drive automation, observability, and incident response while mentoring peers within the STO team.

The role offers flexible work options in Canada—Toronto office, hybrid schedule, or fully remote within Canada—with a preference for GTA candidates who can visit the Toronto office about twice a month.

Qualifications

  • 6+ years in a Site Reliability Engineer role.
  • Experience with cloud platforms (AWS/Azure) and IaC orchestration.

Responsibilities

  • Build scalable, fault-tolerant services across multi-clouds.
  • Define and document runbooks and standard operating procedures.
  • Own and lead STO team delivery projects.
  • Maintain services post-launch with monitoring of availability and latency.
  • Mentor and train other SREs and cross-functional teams.
  • Improve incident response with automated remediation and playbooks.

Skills

SRE fundamentals (SLO/SLI)
Distributed systems
Programming
Cloud-native

Education

Cloud certifications (AWS/Azure)

Tools

Terraform
Ansible
Azure DevOps
Jenkins
Docker
Kubernetes
GitHub

Job description

Senior Site Reliability Developer

Department: SaaS Operations

Employment Type: Full Time

Location: Canada - Flex (0006)

Description

This is a flexible position and has the option of working in our Toronto office full time, hybrid throughout the week or working remotely within Canada.

Preference will be given to candidates located in the GTA who are able to attend the Toronto office approximately two days per month.

Vena is looking for a Senior SRE to join our SaaS Technology and Operations (STO) team. This role is a match for you if you love building highly scalable, resilient and automated services. We are an innovative team which aims to provide exceptional customer experience by leveraging best-in-class automation and orchestration practices for Vena's SaaS platform. As a Senior Site Reliability Developer, you will utilize your software and systems engineering background to build and run large-scale, distributed, fault-tolerant systems and services across AWS and Azure. We strive to hire people who are looking to make an impact and thrive in a flexible work environment driven by business objectives. Your role is to ensure that our systems - both internally and externally facing-have been designed with maximizing resiliency and uptime. Our team focuses on optimizing existing systems, building infrastructure and reducing toil through automation. Practices such as limiting time spent on manual operational work, post-mortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and technical standards.

How You’ll Make an Impact
  • Helping Vena's technology organization build scalable systems, using best practices around automation (reliability) and developer self-service (velocity).
  • Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, planning and reviews. Define and document runbooks and standard operating procedures.
  • You will own and lead projects as part of the STO team’s delivery strategy.
  • Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
  • Provide mentorship and training to other Vena SREs as well as members of the Product and Technology organization on emerging technologies and new processes, drive education and knowledge transfer of design patterns and technical practices.
  • Drive high standards around incident response practices and policies with a focus on automated response and remediation.
  • Participate in influencing and shaping the overall STO team culture.
  • You are a subject-matter expert in one or more technologies leveraged by the Vena platform.
  • Participate in technical interviews for technical positions within STO and occasionally extend into other positions within Vena’s product and technology organization.
  • Participate in on-call rotation.
What we use:

Please note this reflects only a portion of our current technical stack, and we are constantly evolving and revisiting our stack as we grow:

  • A modern multi-cloud infrastructure across AWS and Azure, managed through infrastructure-as-code (Terraform) and configuration-as-code (Ansible)
  • CI/CD primarily through Azure DevOps, with Jenkins still supporting some legacy pipelines
  • Containerized workloads on Amazon EKS, AWS ECS, and Azure Container Apps
  • RDS MySQL, Redshift, Redshift Spectrum, MongoDB, and Elasticsearch
  • Kinesis, SQS, and RabbitMQ (via CloudAMQP)
  • Workflow orchestration with Temporal
  • Identity and access management with Auth0
  • DevOps tools written in Python
  • Back-end applications written using Java, Dropwizard, Spring Boot, and Hibernate
  • Front-end applications written using TypeScript, JavaScript, React (Context Api and Hooks), and Redux
  • Observability and monitoring with Observe Inc. (via OpenTelemetry), CloudWatch, and Azure Monitor
We’d Love to See
  • 6+ years of experience in a Site Reliability Engineer role.
  • You ideally possess an Associate or Professional level certification from AWS or Microsoft Azure.
  • You are adept at core SRE concepts such as SLO, SLIs and error budgets and have direct experience in implementing them.
  • In-depth knowledge of cloud computing platforms (AWS and Azure) and solid experience of setup and management of cloud infrastructure using IaC and orchestration tools.
  • You can write code - in any language. You have implemented your work in a production environment and can back it up with examples.
  • In-depth experience with tools and platforms such as: AWS, Azure, Ansible, Artifact storage (such as Artifactory, ECR), Build/Release Pipelines (such as Azure DevOps, Jenkins, Gitlab, GH Actions, or equivalents), Docker, Github, Kubernetes, Terraform etc.
  • Direct experience with large-scale distributed systems in the cloud using observability and telemetry for oversight of code deployments.
  • Experience with the operational aspects of software systems using telemetry, centralized logging, and alerting with tools such as: CloudWatch, Observe Inc, Prometheus, etc.

The base salary range for this position is $123,250 - 166,750 CAD.

Our salaries are tailored to roles, levels and locations. Your individual pay within this range is influenced by factors like work location, skills, experience and education. As you progress in your role, your compensation may adapt, offering flexibility for growth beyond initial levels. For specifics, your recruiter will provide details and address any questions during the hiring process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Developer
Senior Site Reliability Developer

Eauzone Spa • Toronto

On-site
CAD 123,000 - 167,000
Senior Site Reliability Developer
Senior Site Reliability Developer

Vena Solutions • Toronto

Hybrid
CAD 123,000 - 167,000
Site Reliability Developer 1
Site Reliability Developer 1

Vena Solutions • Toronto

On-site
CAD 89,250 - 120,750
Senior SRE — Cloud Reliability & Automation Lead
Senior SRE — Cloud Reliability & Automation Lead

Vena Solutions • Toronto

Hybrid
CAD 123,000 - 167,000
Senior SRE: Remote/Hybrid Cloud Reliability Leader
Senior SRE: Remote/Hybrid Cloud Reliability Leader

Eauzone Spa • Toronto

Hybrid
CAD 123,000 - 167,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Veerum • Calgary

On-site
CAD 115,000 - 130,000
Flexible benefits
Strong time-off policies
Continuous learning opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Sage Recruiting Inc. • Canada

On-site
CAD 180,000 - 200,000
Unlimited vacation
Comprehensive health and dental benefits
Associate Solution Architect (Full Stack)
Associate Solution Architect (Full Stack)

Vena Solutions • Toronto

On-site
CAD 127,500 - 172,500
Software Architect (Full Stack)
Software Architect (Full Stack)

Vena • Canada

Hybrid
CAD 128,000 - 173,000
Senior or Staff Site Reliability Engineer - 26248
Senior or Staff Site Reliability Engineer - 26248

Enverus • Calgary

On-site
CAD 120,000 - 180,000