Senior Site Reliability Engineer

SimCorp

Toronto

On-site

CAD 100,000 - 141,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Global hybrid work policy
Culture & inclusion
Work-life balance
Career growth

Job summary

SimCorp is hiring a Senior Site Reliability Engineer to own cloud-native services on Azure, focusing on monitoring, observability, cost and compliance. You will partner with DevOps, clients, and stakeholders to ensure reliability, performance, and automation in Azure environments, while onboarding long-running clients and driving operational excellence.

Ideal candidates have 3+ years in SRE/DevOps, deep Azure expertise, and strong IaC, monitoring and security skills.

Qualifications

  • Bachelor’s degree in Computer Science or related field (Master’s is a plus).
  • 3+ years in Site Reliability, DevOps, or Cloud Engineering roles.
  • Must have expertise with Microsoft Azure Cloud.
  • Expertise in IaC using Bicep, ARM and Terraform.
  • Solid experience in monitoring and logging tools (Azure Monitor, Application Insights, DataDog, Log Analytics).
  • Hands-on experience in IdP Onboarding and integrating IdP solutions like Azure Entra ID, Okta, KeyCloak or PingFederate.
  • Experience in centralizing authentication, managing user identities, and implementing secure access protocols (SAML, OAuth, OIDC).
  • Experience with observability frameworks like Open Telemetry and distributed tracing systems.
  • Experience with application reliability platforms like Checkly or equivalent.
  • Experience setting up synthetic monitoring using Playwright or equivalent.
  • Knowledge of AI/ML-based anomaly detection, log aggregation and analysis tools like Microsoft Azure Anomaly Detector or equivalent.
  • Experience with Microsoft Defender Suite (EDR, XDR) and Sentinel.
  • Proficient in KQL for threat hunting and improving compliance scores using Defender for Cloud.
  • Able to identify and remediate vulnerabilities.
  • Understanding of networking, containerization (Kubernetes, Docker).
  • Good understanding of APIs, scripting languages like PowerShell, Bash, Kusto and databases like SQL, Cosmos DB and Postgres SQL
  • Familiarity with SimCorp Dimension & Sales force is a plus
  • Proficiency in IT service management (ITSM) frameworks like ITIL, focusing on incident, change, and problem management to improve operational efficiency
  • Experience managing both onboarding projects and live production operations
  • Collaborative mindset and ability to work in cross-functional teams
  • Interest in continuous learning and growth within your Product Area

Responsibilities

  • Support the operational and enhancement of mission-critical environments for both new and existing Cloud Native products & services
  • Collaborate with product development teams to enhance monitoring, observability, reliability, and performance of these services.
  • Collaborate deeply across engineering teams to understand systems at the code level.
  • Manage & improve our infrastructure deployment pipelines and troubleshoot onboarding and operational issues
  • Drive capacity planning efforts to ensure our platform is resilient and scalable as we grow.
  • Build tools and automation to eliminate manual TOIL, improve engineering velocity, developer experience, and improve system reliability.
  • Define and manage SLOs and error budgets in partnership with Engineering teams.
  • Contribute to incidents, problems, and change management processes.
  • Execute disaster recovery, configuration management, and platform readiness tasks.
  • Flexible working in regular & evening shift on rotational basis and provide weekend or On-Call support as needed.
  • Collaborate with Agile teams and take part in design discussions with clients, vendors, and stakeholders.
  • Contribute to knowledge sharing across multiple Product Areas.
  • Leverage a strong foundation in ITIL practices, including problem, change, and incident management.

Skills

Azure Cloud
Kubernetes
Docker
PowerShell
Bash
Kusto
SQL
Cosmos DB
PostgreSQL
ITIL

Education

Bachelor’s degree in Computer Science or related field

Tools

Terraform
Bicep
ARM templates
Azure Monitor
DataDog
Log Analytics
Okta
KeyCloak
PingFederate
OpenTelemetry

Job description

WHAT MAKES US, US

Join some of the most innovative thinkers in FinTech as we lead the evolution of financial technology. If you are an innovative, curious, collaborative person who embraces challenges and wants to grow, learn and pursue outcomes with our prestigious financial clients, say Hello to SimCorp! At its foundation, SimCorp is guided by our values – caring, customer success-driven, collaborative, curious, and courageous. Our people-centered organization focuses on skills development, relationship building, and client success. We take pride in cultivating an environment where all team members can grow, feel heard, valued, and empowered. If you like what we’re saying, keep reading!

WHY THIS ROLE IS IMPORTANT TO US

As a Senior Site Reliability Engineer, you will be working on Cloud Native Products & Services, taking ownership of various responsibility domains like monitoring, observability, release management, vulnerability management, cost management, audit & compliance etc. You will work closely with DevOps engineers, clients, and stakeholders to ensure reliability, performance, and automation for both existing and new cloud native products & services. Onboard and long-running clients on them. Your contributions will drive stability, continuous improvement, and operational excellence in our Azure-based environments. This role blends hands-on engineering, incident response, platform configuration, and service quality, - guided by ITIL and SRE best practices.

WHAT YOU WILL BE RESPONSIBLE FOR
  • Support the operational and enhancement of mission-critical environments for both new and existing Cloud Native products & services
  • Collaborate with product development teams to enhance monitoring, observability, reliability, and performance of these services.
  • Collaborate deeply across engineering teams to understand systems at the code level.
  • Manage & improve our infrastructure deployment pipelines and troubleshoot onboarding and operational issues
  • Drive capacity planning efforts to ensure our platform is resilient and scalable as we grow.
  • Build tools and automation to eliminate manual TOIL, improve engineering velocity, developer experience, and improve system reliability.
  • Define and manage SLOs and error budgets in partnership with Engineering teams.
  • Contribute to incidents, problems, and change management processes.
  • Execute disaster recovery, configuration management, and platform readiness tasks.
  • Flexible working in regular & evening shift on rotational basis and provide weekend or On-Call support as needed.
  • Collaborate with Agile teams and take part in design discussions with clients, vendors, and stakeholders.
  • Contribute to knowledge sharing across multiple Product Areas.
  • Leverage a strong foundation in ITIL practices, including problem, change, and incident management.
WHAT WE VALUE
  • Bachelor’s degree in Computer Science or related field (Master’s is a plus)
  • 3+ years in Site Reliability, DevOps, or Cloud Engineering roles
  • Must have expertise with Microsoft Azure Cloud.
  • Expertise in Infrastructure as Code (IaC) using Bicep, ARM and Terraform.
  • Solid experience in monitoring and logging tools (Azure Monitor, Application Insights, DataDog, Log Analytics).
  • Hand-on experience in IdP Onboarding and integrating, configuring IdP solutions like Azure Entra ID, Okta, KeyCloak or PingFederate.
  • Experience in centralizing authentication, managing user identities, and implementing secure access protocols (SAML, OAuth, OIDC)
  • Experience working with observability frameworks like Open Telemetry and distributed tracing systems
  • Experience working with application reliability platforms like Checkly or equivalent
  • Experience setting up synthetic monitoring using Playwright or equivalent
  • Knowledge of AI/ML-based anomaly detection, log aggregation and analysis tools like Microsoft Azure Anomaly Detector or equivalent
  • Experience working with Microsoft Defender Suite (EDR, XDR) and Sentinel.
  • Proficient in KQL for threat hunting and improving compliance scores using Defender for Cloud.
  • Able to identify and remediate vulnerabilities
  • Understanding of networking, containerization (Kubernetes, Docker)
  • Good understanding of APIs, scripting languages like PowerShell, Bash, Kusto and databases like SQL, Cosmos DB and Postgres SQL
  • Familiarity with SimCorp Dimension & Sales force is a plus
  • Proficiency in IT service management (ITSM) frameworks like ITIL, focusing on incident, change, and problem management to improve operational efficiency
  • Experience managing both onboarding projects and live production operations
  • Collaborative mindset and ability to work in cross-functional teams
  • Interest in continuous learning and growth within your Product Area
Benefits
  • Global hybrid work policy - We ask you to work 2 days a week from the office. If you choose you can work remotely the other days. Of course, you are welcome at the office if that is your preference.
  • Culture – Inclusive and diverse company culture
  • Work-life balance – We believe that an equilibrium between professional responsibilities makes us all the best version of ourselves, both in private life and as colleagues in the workplace
  • Empowerment – We believe that all voices are valuable and must be heard. You will be involved in shaping our work processes
  • Career & Growth – Simcorp does offer opportunities for professional development: there is never just only one route - we offer an individual approach to professional development to support the direction you want to take.
WHO WE ARE

For over 50 years, we have worked closely with investment and asset managers to become the world’s leading provider of integrated investment management solutions. We are 3,000+ colleagues with a broad range of nationalities, education, professional experiences, ages, and backgrounds. SimCorp is an independent subsidiary of the Deutsche Börse Group. Following the recent merger with Axioma, we leverage the combined strength of our brands to provide an industry-leading, full, front-to-back offering for our clients. SimCorp is an equal opportunity employer and welcome applicants from all backgrounds, without regard to race, gender, age, disability, or any other protected status under applicable law. We are committed to building a culture where diverse perspectives and expertise are integrated into our everyday work. We believe in the continual growth and development of our employees, so that we can provide best-in-class solutions to our clients. For Toronto City only: The annual base salary range for this position is 100 000,00 - 140 600,00 CAD. Additionally, employees are eligible for an annual discretionary bonus, and benefits including health care, leave, and retirement plans. Your total compensation may vary based on role, location, department and individual performance. #Li-Hybrid SimCorp is a provider of industry-leading integrated investment management solutions for the global buy side. Founded in 1971, with more than 3,000 employees across five continents, SimCorp is a truly global technology leader that empowers more than half of the world’s top 100 financial companies through its integrated platform, services, and partner ecosystem. SimCorp is a subsidiary of Deutsche Börse Group. As of 2024, SimCorp includes Axioma, the leading provider of risk and management and portfolio optimization solutions for the global buy side. ---------------------------------------------------------------- Discover more about our culture, recruitment process, and commitment to promoting meaningful work

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

SimCorp • Toronto

On-site
CAD 113,000 - 142,000
Health and dental care
Group RRSP/TFSA
Hybrid work policy
Senior DevOps Engineer
Senior DevOps Engineer

SimCorp • Toronto

On-site
CAD 98,000 - 142,000
Annual discretionary bonus
Health and dental care
Time off
+1
Lead Site Reliability Engineer - Interface & Connectivity
Lead Site Reliability Engineer - Interface & Connectivity

SimCorp • Toronto

On-site
CAD 140,000 - 180,000
Hybrid work policy
Global benefits
Senior Site Reliability Engineer - Batch Management
Senior Site Reliability Engineer - Batch Management

SimCorp • Toronto

On-site
CAD 100,000 - 150,000
Hybrid work
Pension
IP sprints
+2
Principal/Senior Principal Investment Accounting Solution Consultant – IAS Value Team
Principal/Senior Principal Investment Accounting Solution Consultant – IAS Value Team

SimCorp • Toronto

On-site
CAD 199,000 - 250,000
Salary and bonus schemes
Work-life balance
Professional development
Senior Cyber Defense Engineer
Senior Cyber Defense Engineer

SimCorp • Toronto

Hybrid
CAD 114,000 - 171,000
Global hybrid policy
Remote work available part-time
Senior Site Reliability Engineer - Batch Management
Senior Site Reliability Engineer - Batch Management

Socket.dev • Toronto

Hybrid
CAD 90,000 - 130,000
Hybrid work schedule
Professional development opportunities
IP sprints and skill growth
Lead Business Consultant, Alternative Investments and Private Credit
Lead Business Consultant, Alternative Investments and Private Credit

SimCorp • Toronto

On-site
CAD 119,000 - 148,000
Attractive salary
Bonus scheme
Pension
+2
Senior Financial Data Analyst
Senior Financial Data Analyst

SimCorp • Toronto

On-site
CAD 101,000 - 140,000
Health and dental care
Annual discretionary bonus
Group RRSP/TFSA
Lead Business Consultant, IBOR
Lead Business Consultant, IBOR

SimCorp • Toronto

On-site
CAD 115,000 - 165,000
Salary + bonus scheme
Pension
Hybrid work model
+1