Site Reliability Engineer (1890)

Direct IT Recruiting Inc.

Toronto

Hybrid

CAD 117,000 - 129,000

Full time

9 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Direct IT Recruiting Inc. is seeking a hands-on Site Reliability Engineer in a hybrid Toronto office role. The contract runs for 9 months, with 8 hours per day and 40 hours per week, and includes a 3-day on-site schedule.

The role emphasizes Python development, automation, CI/CD, observability, and platform modernization. You will work with GitHub Actions, Argo CD, Dynatrace, ELK, and related tooling to improve health checks and deployment pipelines.

Qualifications

  • 6+ years in Site Reliability Engineering, production engineering, or related tech services.
  • Strong Python development and scripting with good testing and documentation.

Responsibilities

  • Build automation and scripting solutions using Python to reduce manual efforts.
  • Improve CI/CD and deployment processes with GitHub Actions and Argo CD.
  • Advance SRE practices to improve availability, performance, and resilience of systems.

Skills

Python
GitHub Actions
CI/CD
Automation
Observability
Dynatrace
Argo CD
PowerShell
ELK

Tools

Git
JIRA
Confluence

Job description

NOTE: Hybrid 3 days a week in Toronto Office

Type: 9 month contract, 8 hours/day, 40 hours/week

Rate: $85 to $94/hr

SKILLS: Python, GitHub Actions, CI/CD, automation, Site Reliability Engineering, Dynatrace, Argo CD, PowerShell, observability

Industry: Capital Markets and Investment Management

DESCRIPTION

We are seeking an experienced Site Reliability Engineer to support applications, data pipelines, platforms, and services across Capital Markets, Investment Analytics, and Total Fund Management.

This is a hands on individual contributor role focused on Python development, operational automation, CI/CD, observability, disaster recovery readiness, and platform modernization. The successful candidate will improve health checks, deployment pipelines, telemetry, diagnostics, and operational processes while addressing a significant backlog of reliability initiatives.

The environment is highly Python focused and includes GitHub Actions, Argo CD, Dynatrace, Elastic and ELK, PowerShell, Ansible, Bash, and cron jobs. Incident volume is relatively low, allowing the successful candidate to focus primarily on proactive engineering and sustainable automation.

Capital markets experience is beneficial but not required. Strong Python, automation, CI/CD, and observability experience are the primary priorities.

RESPONSIBILITIES
  • Build Automation and Scripting Solutions: Develop maintainable Python scripts, tools, and automation solutions that reduce manual effort, improve consistency, and strengthen operational efficiency.
  • Improve CI/CD and Deployment Processes: Enhance build, deployment, release, rollback, and production validation processes using GitHub Actions, Argo CD, and related technologies.
  • Advance Site Reliability Engineering Practices: Improve the availability, performance, resilience, supportability, and operational sustainability of applications, data pipelines, platforms, and services.
  • Enhance Observability and Telemetry: Strengthen application health checks, logging, monitoring, metrics, alerts, dashboards, diagnostics, and telemetry. Support the adoption of Dynatrace and the transition from the current Elastic and ELK logging environment.
  • Strengthen Incident Response and Root Cause Analysis: Investigate production issues using logs, metrics, traces, and diagnostic tools. Identify root causes and implement corrective and preventative actions.
  • Improve Detection and Recovery: Develop monitoring, diagnostics, automation, and operational runbooks that reduce Mean Time to Detect and Mean Time to Restore.
  • Improve Operational Readiness: Conduct production readiness reviews and ensure monitoring, runbooks, resilience testing, documentation, release processes, and support arrangements are in place before production deployment.
  • Support Disaster Recovery and Resilience: Contribute to disaster recovery plans, recovery procedures, resilience initiatives, and operational testing.
  • Support Production Systems: Participate in incident response, coordinate technical activities, communicate service impacts, and elevate major or cross service incidents through the appropriate channels.
  • Partner Across Technology Teams: Collaborate with Product Engineering, Data Solutions, Data Platform, Architecture, Security, Technology Services, database, and quality assurance teams to identify and resolve reliability, resilience, and supportability risks.
REQUIREMENTS
  • 6 or more years of experience in Site Reliability Engineering, production engineering, platform engineering, application support, or technology service delivery.
  • Strong hands on Python development and scripting experience, including sound coding, testing, documentation, and source control practices.
  • Experience designing automation solutions and independently selecting the appropriate scripting language, platform, or tool to solve operational problems.
  • Experience developing, maintaining, or improving CI/CD pipelines, preferably using GitHub Actions and Argo CD.
  • Experience with observability, application monitoring, centralized logging, telemetry, alerting, and production diagnostics. Dynatrace experience is strongly preferred.
  • Experience with Elastic, ELK, PowerShell, Ansible, Bash, cron jobs, or similar scripting and automation technologies is beneficial.
  • Hands on experience operating and improving highly available applications, data pipelines, platforms, APIs, message queues, or distributed services.
  • Experience with incident response, root cause analysis, disaster recovery, resilience testing, capacity planning, and production readiness.
  • Experience with cloud environments, infrastructure as code, DevOps, DataOps, automated testing, release automation, and production validation practices.
  • Experience using Git, JIRA, Confluence, and related engineering collaboration tools.
  • Strong problem solving skills with the ability to understand an unfamiliar problem space, establish an effective approach, and independently drive work through completion.
  • Strong communication and collaboration skills, including the ability to explain technical risks, incidents, and service impacts in clear business language.
  • Experience working with developers, delivery leads, architects, business analysts, database administrators, quality assurance professionals, security teams, and infrastructure teams.
  • Knowledge of Agile, Waterfall, DevOps, ITIL, or COBIT practices.
  • Practical experience using approved AI assisted engineering tools to troubleshoot applications, databases, infrastructure, code, logs, and telemetry is an asset.
  • Capital markets, investment analytics, investment management, or financial services experience is beneficial but not required.

Note: As part of our hiring process, we use AI based systems to support initial applicant screening.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — Python Automation & CI/CD
Site Reliability Engineer — Python Automation & CI/CD

Direct IT Recruiting Inc. • Toronto

Hybrid
CAD 117,000 - 129,000
Site Reliability Engineer
Site Reliability Engineer

Kyndryl • Toronto

On-site
CAD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Ian Martin Group • Toronto

On-site
CAD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

ALLTECH CONSULTING SVC INC • Quebec

On-site
CAD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

ApTask • Montreal

On-site
CAD 125,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

Rippling, Inc. • Toronto

On-site
CAD 110,000 - 165,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Akkodis • Toronto

On-site
CAD 140,000 - 165,000
Bonus
Benefits
Senior QA Automation Engineer
Senior QA Automation Engineer

Global Technical Talent, an Inc. 5000 Company • Toronto

Hybrid
CAD 124,000 - 138,000
Medical Insurance
401k Retirement Fund
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 130,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Akkodis • Toronto

Hybrid
CAD 120,000 - 180,000