Manager, Site Reliability Engineering

MasterCard

Vancouver

On-site

CAD 121,000 - 169,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mastercard Vancouver seeks a Manager, Site Reliability Engineering to lead a team responsible for the reliability, scalability, and performance of our global applications. You will drive architecture decisions, mentor engineers, and champion automated operations.

In this role you will set monitoring, incident response, and capacity planning standards, partner with development and product teams to enable zero-downtime deployments, and continually improve system health across Mastercard's services

Qualifications

  • BS degree in Computer Science or related field.
  • Coding or scripting experience.
  • Interest in automation and scalable systems.
  • Experience with distributed systems and DevOps practices.

Responsibilities

  • Serve as primary contact for application health and capacity.
  • Lead automation and incident response initiatives.
  • Collaborate with cross-functional teams on scalability strategies.
  • Oversee incident response and post-mortem processes.
  • Mentor engineers and drive operational excellence.

Skills

SRE leadership
Automation mindset
Problem solving
Cross-functional collaboration

Education

BS degree in Computer Science or related field

Tools

Chef
Ansible
Artifactory
GitHub
Jenkins
XLR
Remedy

Job description

Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Site Reliability Engineering

Who is Mastercard?

At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits everyone, everywhere, by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team – one that makes better decisions, drives innovation, and delivers better business results.

Technology at Mastercard

What we create today will define tomorrow. Revolutionary technologies that reshape the digital economy to be more connected and inclusive than ever before. Safer, faster, more sustainable and we need the best people to do it. Technologists who are energized by the challenges of a truly global network. With the talent and vision to create the critical systems and products that power global commerce and connect people everywhere to the vital goods and services they need every day.

About the Role

The Business Operations team is seeking a highly motivated and experienced Manager, Site Reliability Engineering (SRE) to join our team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard's global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor.

The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for Mastercard products. As Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture.

As part of the Business Operations team, you will:
  • Oversee a team of individual contributors, supporting the execution of strategic initiatives by providing technical expertise and leadership within the Site Reliability Engineering discipline to analyze complex problems and provide novel solutions and/or improvements.
  • Guide the team in automating routine tasks, troubleshooting complex issues, and optimizing system performance.
  • Collaborate with cross-functional teams to develop strategies for system scalability and resilience, training team members on technical skills, operational best practices, and incident management.
  • Oversee incident response efforts, ensuring timely resolution and comprehensive root cause analysis.
  • Cultivate a culture of continuous improvement by promoting best practices, innovation, and proactive risk management.
  • Support the implementation and maintenance of high-availability systems to ensure operational stability.
  • Contribute to documentation, knowledge sharing, and best practices to improve team operational procedures.
  • Lead automation and scripting efforts to streamline operational processes and incident response workflows.
  • Manage a team of individual contributors(s) and/or technical lead(s), directing area processes and work to ensure that they align with functional best practices and organizational standards; conduct goal setting and performance appraisal processes to coach team members and support their professional development.
Role:
  • Serve as the primary contact responsible for the overall application health, performance, and capacity
  • Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
  • Partner with the development and product team of a new application to establish the right monitoring and alerting strategy and create the framework to achieve zero downtime during deployment.
  • Serve as the primary contact responsible for ensuring application scalability, performance, and resilience.
  • Practice sustainable incident response and blameless post-mortems while taking a holistic approach to problem-solving and optimizing time to recover.
  • Automate data-driven alerts to proactively elevate issues.
  • Work with development teams to establish SLOs and improve reliability.
  • Tackle complex development, automation, and business process problems.
  • Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation, and refinement.
  • Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead Mastercard in DevOps automation and best practices.
  • Increase automation and tooling to reduce toil and manual interventiono Analyses ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns
All about you:
  • BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics), or equivalent practical experience.
  • Coding or scripting exposure.
  • Appetite for change and pushing the boundaries of what can be done with automation. Be curious about new technology, infrastructure, and practices to scale our architecture and prepare for future growth.
  • Experience with algorithms, data structures, scripting, pipeline management, and software design
  • Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive.
  • Interest in designing, analyzing, and troubleshooting large-scale distributed systems.
  • Willingness and ability to learn and take on challenging opportunities and to work as a member of a matrix-based, diverse and geographically distributed project team.
  • Ability to balance doing things right with fixing things quickly. Flexible and pragmatic, while working towards improving the long-term health of the system.
  • Comfortable collaborating with cross-functional teams to ensure that expected system behaviour is understood and monitoring exists to detect anomalies.
  • Experience with DevOps practices and tools such as Chef, Ansible, Artifactory, GitHub, Bitbucket, Jenkins, XLR, and Remedy.

Mastercard is a merit-based, inclusive, equal opportunity employer that considers applicants without regard to gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. We hire the most qualified candidate for the role.

Corporate Security Responsibility

All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must: Abide by Mastercard’s security policies and practices; Ensure the confidentiality and integrity of the information being accessed; Report any suspected information security violation or breach, and Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.

In line with Mastercard’s total compensation philosophy and assuming that the job will be performed in Canada, the successful candidate will be offered a competitive pay based on location, experience and other qualifications for the role and may be eligible to participate in a discretionary annual incentive program.

Pay Ranges Vancouver, Canada: $121,000 - $169,000 CAD

Everyone wants easier ways to pay; we invent them. Checkout lines are slow; we speed them along. Merchants want more sales; we give them data and insights. People need financial access; we connect them. Corporate purchasing is complicated; we make it simple. Commuters are busy; we speed them on their way. Governments need greater efficiencies; we help create them. Small businesses are virtual; we give them access to a world of buyers. Retailers want to fight fraud; we provide the tools.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, Software Engineering
Manager, Software Engineering

MasterCard • Toronto

On-site
CAD 127,000 - 203,000
Lead Software Engineer (SDET)
Lead Software Engineer (SDET)

MasterCard • Toronto

On-site
CAD 127,000 - 203,000
Manager, Product Management
Manager, Product Management

MasterCard • Toronto

On-site
CAD 121,000 - 169,000
Manager, Software Engineering
Manager, Software Engineering

Mastercard • Vancouver

On-site
CAD 127,000 - 203,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

Mastercard • Toronto

On-site
CAD 121,000 - 169,000
Senior Software Engineer (Java)
Senior Software Engineer (Java)

Mastercard • Vancouver

On-site
CAD 111,000 - 160,000
Associate Managing Consultant, Advisors & Consulting Services, Marketing
Associate Managing Consultant, Advisors & Consulting Services, Marketing

MasterCard • Toronto

On-site
CAD 83,000 - 132,000
Associate Managing Consultant, Advisors & Consulting Services, Deploy
Associate Managing Consultant, Advisors & Consulting Services, Deploy

Mastercard • Toronto

On-site
Associate Consultant, Launch Program 2027 - Vancouver, Canada
Associate Consultant, Launch Program 2027 - Vancouver, Canada

MasterCard • Vancouver

On-site
CAD 79,000 - 96,000
Senior Solutions / Sales Engineer
Senior Solutions / Sales Engineer

Mastercard • Toronto

On-site
CAD 83,000 - 132,000