Site Reliability Engineer

Gemini Solutions Pvt Ltd

Toronto

On-site

CAD 120,000 - 170,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Gemini Solutions Pvt Ltd seeks a Senior Site Reliability Engineer to drive reliability, scalability, and performance for distributed systems. You will bridge software engineering, cloud infrastructure, and production operations while leading platform reliability initiatives.

You will own availability, define SLIs/SLOs, and lead incident response across global environments. Strong automation, IaC, and cloud experience are essential for long-term resilience and scalable growth.

Qualifications

  • 3+ years in SRE/DevOps/Production Engineering.
  • Experience in production-critical environments with high availability requirements.
  • Exposure to global systems and cross-team collaboration.

Responsibilities

  • Own availability, performance, and scalability of production systems.
  • Define and implement SLIs, SLOs, and error budgets.
  • Drive continuous improvements in system resilience and efficiency.
  • Lead end-to-end incident response and service restoration.
  • Perform root cause analysis across infrastructure, application, data, and network layers.
  • Implement long-term fixes to reduce recurrence.
  • Design and enhance monitoring, logging, and alerting systems.
  • Develop dashboards and improve alert quality.
  • Enable proactive detection of system issues.
  • Automate operational workflows and build CI/CD pipelines.
  • Implement IaC for scalable infrastructure management.
  • Manage and optimize systems on modern cloud platforms.
  • Troubleshoot distributed systems across compute, storage, and network layers.
  • Diagnose latency, routing, and performance issues in globally distributed environments.
  • Troubleshoot data pipelines, job failures, and data inconsistencies.
  • Perform data validation and analysis.
  • Ensure reliability across data dependencies and workflows.
  • Diagnose DNS, HTTP/S, proxies, and load balancing issues.
  • Work with CDN/edge platforms to optimize traffic routing and performance.
  • Communicate system status, incidents, and risks with stakeholders.
  • Partner with cross-functional teams to drive reliability improvements.
  • Apply AI/ML-driven techniques for anomaly detection and alert optimization.

Skills

Python automation
Bash scripting
Cloud platforms
Observability tools
Docker & Kubernetes
Networking & CDN
CI/CD pipelines
IaC (Terraform/Ansible)
Datadog/Splunk/Prometheus/Grafana
AI-driven reliability

Tools

Jenkins
GitLab CI
Terraform
CloudFormation

Job description

  • We are looking for a Senior Site Reliability Engineer (SRE) with a strong platform
  • ownership mindset to drive reliability, scalability, and performance of mission-critical, distributed systems.
  • This role sits at the intersection of software engineering, cloud infrastructure, and production operations, with a focus on building resilient systems, improving observability, automating operations, and driving reliability at scale.
  • You will act as a technical lead for platform reliability, working closely with engineering and business stakeholders to ensure systems are highly available, performant, and continuously improving.
Experience:
  • 3+ years of experience in SRE, DevOps, or Production Engineering
  • Experience working in production-critical environments with high availability requirements
  • Exposure to global systems and cross-team collaboration
Key Responsibilities
Platform Reliability & Ownership
  • Own availability, performance, and scalability of production systems
  • Define and implement SLIs, SLOs, and error budgets
  • Drive continuous improvements in system resilience and efficiency
Incident Management & Root Cause Analysis
  • Lead end-to-end incident response and service restoration
  • Perform deep root cause analysis across infrastructure, application, data, and network layers
  • Implement long-term fixes and reduce recurrence through engineering improvements
Observability & Monitoring
  • Design and enhance monitoring, logging, and alerting systems
  • Develop actionable dashboards and improve alert quality
  • Enable proactive detection of system issues
Automation & DevOps Practices
  • Automate operational workflows to reduce manual effort
  • Build and maintain CI/CD pipelines
  • Implement Infrastructure as Code (IaC) for scalable infrastructure management
  • Manage and optimize systems on modern cloud platforms
  • Troubleshoot distributed systems across compute, storage, and network layers
  • Diagnose latency, routing, and performance issues in globally distributed environments
Data & Workflow Reliability
  • Troubleshoot data pipelines, job failures, and data inconsistencies
  • Perform data validation and analysis
  • Ensure reliability across data dependencies and workflows
Networking & Traffic Management
  • Diagnose issues related to DNS, HTTP/S, proxies, and load balancing
  • Work with CDN and edge delivery platforms (e.g., Akamai or similar) to optimize traffic routing and performance
Stakeholder Collaboration
  • Act as a liaison between engineering teams and business stakeholders
  • Communicate system status, incidents, and risks with clarity and context
  • Partner with cross-functional teams to drive reliability improvements
AI-Driven Reliability (Emerging Focus)
  • Apply AI/ML-driven techniques for anomaly detection, alert optimization, and
  • predictive issue identification
  • Leverage intelligent automation to improve incident response and operational
Core Expectations
  • Demonstrates strong ownership of production systems and outcomes
  • Independently drives incident resolution and follow-through
  • Applies structured, analytical thinking to complex technical problems
  • Communicates effectively in high-impact, production-critical scenarios
  • Focuses on long-term reliability and scalability improvements
Technical Skills:
Programming & Automation
  • Strong experience in Python for automation and tooling
  • Proficiency in shell scripting (Bash)
  • Experience with API-driven and event-driven automation
  • Hands-on experience with AWS, Azure, or GCP
  • Strong understanding of cloud architecture, networking, and security fundamentals
  • Infrastructure as Code using Terraform, CloudFormation, or Ansible
DevOps & CI/CD
  • Experience with Jenkins, GitLab CI, or similar tools
  • Strong understanding of build, release, and deployment pipelines
Observability
  • Experience with Datadog, Splunk, Prometheus, or Grafana
  • Strong logging, monitoring, and alerting practices
  • Familiarity with incident management tools (e.g., PagerDuty)
Data & Databases
  • Strong SQL skills for troubleshooting and validation
  • Understanding of data pipelines and system dependencies
Systems & Platform
  • Experience with Docker and containerized environments
  • Exposure to Kubernetes and web servers (e.g., Nginx)
Orchestration
  • Experience with Airflow, Autosys, or similar scheduling tools
Networking & CDN
  • Strong understanding of DNS, HTTP/S, proxies, and load balancing

Experience with CDN and edge delivery platforms (e.g., Akamai or similar)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Mantu • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Software Development Lead- Platform
Software Development Lead- Platform

IMDEX • Calgary

On-site
CAD 140,000 - 200,000
Senior Software Engineer, Platform (SRE)
Senior Software Engineer, Platform (SRE)

HRB • Kitchener, Southwestern Ontario

On-site
CAD 120,000 - 180,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind Americas • Montreal (administrative region)

On-site
CAD 110,000 - 170,000
Competitive salary
Laptop provided
Professional development
+2
Site Reliability Engineer
Site Reliability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 90,000 - 130,000
Staff Site Reliability Engineer - Confluent Incident Management & Reliability
Staff Site Reliability Engineer - Confluent Incident Management & Reliability

IBM • Toronto

On-site
CAD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Montreal (administrative region)

Hybrid
CAD 90,000 - 130,000
Site Reliability Engineer (Linux / Cloud Infrastructure)
Site Reliability Engineer (Linux / Cloud Infrastructure)

Atlantis IT Group • Montreal

On-site
CAD 80,000 - 100,000
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes
[8SN] Senior Site Reliability Engineer (SRE) – Kubernetes

Worky • Montreal (administrative region)

On-site
CAD 120,000 - 170,000
Laptop
Flexible work arrangements
Professional development and training
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cerebras • Vancouver

On-site
CAD 100,000 - 130,000