Site Reliability Engineer

Kyndryl

Toronto

Hybrid

CAD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A global technology services provider is seeking a Site Reliability Engineer in Toronto to enhance the reliability and efficiency of critical batch workloads. This mid-senior level contract role emphasizes automation, application development, and observability using Dynatrace. Candidates should possess expert skills in Python, Linux systems engineering, and experience with CI/CD tools. This position offers a hybrid work model requiring 2-3 days onsite and is not a full-time role with the company.

Qualifications

  • Expert-level Python skills, including performance tuning and testing.
  • Strong Linux systems engineering expertise (kernel tuning, networking).
  • Proven experience optimizing batch workloads for performance.

Responsibilities

  • Engineer resilient batch processing pipelines to reduce runtime.
  • Implement Dynatrace dashboards and maintain system health visibility.
  • Configure Linux/Windows environments for optimal reliability.

Skills

Python programming
Linux systems engineering
Optimizing batch workloads
Dynatrace observability
Apache Airflow
Distributed systems concepts
CI/CD pipelines
Infrastructure as Code
Containers and orchestration
Incident management

Tools

GitHub Actions
Azure DevOps
Jenkins
Terraform
Ansible
Docker
Kubernetes

Job description

Join to apply for the Site Reliability Engineer role at Kyndryl.

Direct message the job poster from Kyndryl.

Recruitment & Strategic Staffing @Kyndryl | Partnering with IT Consultants in Financial Services & Technology
  • Position: Site Reliability Engineer
  • Client: Financial Services - Capital Markets Technology
  • Duration: 12-month contract with potential extensions
  • Location: Toronto, Canada - 2 to 3 days onsite per week
  • Language: English
  • Hours: 37.5 hours/week

Our client is looking for a Site Reliability Engineer (SRE) to enhance the reliability, performance, and efficiency of mission‑critical batch workloads across Capital Markets Technology. The SRE will serve as a technical lead focused on automation, application development, systems performance engineering, and observability using Dynatrace. This position is pivotal in driving operational excellence and maturing reliability practices across the organization.

Qualifications
  • Expert‑level Python skills, including performance tuning, concurrency (async/multiprocessing), testing, and packaging.
  • Strong Linux systems engineering expertise (kernel tuning, networking, process management, filesystem optimization).
  • Proven experience optimizing batch workloads for performance, reliability, and cost efficiency.
  • Deep knowledge of Dynatrace for observability (dashboards, KPIs, tagging, alerts, anomaly detection).
  • Hands‑on experience with Apache Airflow (DAG design, scheduler tuning, SLA management).
  • Strong understanding of distributed systems concepts — retries, idempotency, backpressure, data integrity.
  • Experience with CI/CD pipelines (GitHub Actions, Azure DevOps, Jenkins) and Infrastructure as Code (Terraform, Ansible).
  • Familiarity with containers and orchestration tools (Docker, Kubernetes).
  • Excellent incident management, troubleshooting, and communication skills.
Responsibilities
  • Reliability & Performance: Engineer resilient and performant batch processing pipelines by reducing runtime and minimizing failures.
  • Observability: Implement and maintain Dynatrace dashboards, alerts, and runbooks to ensure deep visibility into system health.
  • Systems Engineering: Configure and tune Linux and Windows environments for optimal reliability and speed.
  • Automation & Orchestration: Design and refine Airflow DAGs, automate deployments with CI/CD pipelines, and reduce operational toil through code.
  • Incident Management: Lead incident response, conduct root‑cause analysis, and implement improvements based on post‑mortems and SLOs.
  • Security & Compliance: Ensure all reliability and automation processes adhere to security best practices and regulatory compliance standards.

Please note this is for a contract position with one of our clients and not a full-time employment role with Kyndryl Canada.

Seniority level
  • Mid‑Senior level
Employment type
  • Contract
Job function
  • Information Technology
Industries
  • IT Services and IT Consulting

Referrals increase your chances of interviewing at Kyndryl by 2x.

Sign in to set job alerts for “Site Reliability Engineer” roles.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Applications System Administrator
Software Engineering Applications System Administrator

Kyndryl • Ottawa

Hybrid
CAD 85,000 - 110,000
Service Delivery Manager
Service Delivery Manager

Kyndryl • Ottawa

On-site
CAD 80,000 - 110,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Orion Innovation • Quebec

On-site
CAD 90,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

RXinsider LTD. • Montreal (administrative region)

Hybrid
CAD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

ApTask • Montreal

On-site
CAD 125,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Thinkific • Canada

Remote
CAD 111,000 - 167,000
Fair and transparent pay
Inclusive work culture
Remote work flexibility
IT Project Manager
IT Project Manager

Kyndryl • Toronto

Hybrid
CAD 90,000 - 130,000
DevOps Engineer/ Dynatrace
DevOps Engineer/ Dynatrace

Motion Recruitment • Toronto

On-site
CAD 110,000 - 140,000
Medical, Dental, and Vision Insurance
Vacation Time
Client Partner
Client Partner

Kyndryl • Regina

On-site
CAD 90,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Tecsys Inc. • Montreal (administrative region)

On-site
CAD 90,000 - 120,000
Digital-first work environment
Collaborative workspaces
Continuous learning opportunities