Kubernetes Reliability Engineer

Roche Holding AG

Mississauga

On-site

CAD 106,000 - 139,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Hoffmann-La Roche Ltée. au Canada recherche un Kubernetes Reliability Engineer pour concevoir, déployer et optimiser une plateforme Kubernetes sur des environnements hybrides (on‑prem et AWS). Vous appliquerez des principes logiciels à l'exploitation, avec un focus sur l'automatisation et l'observabilité.

Le candidat idéal maîtrise Kubernetes, IaC (Ansible/Terraform) et comprend l'UX des pipelines CI/CD, avec une expérience démontrée dans des environnements critiques et régulés.

Qualifications

  • Diplôme de Bachelor en informatique, mathématiques, physique ou domaine connexe et 2–5 ans d'expérience pertinente.
  • Sans diplôme: 4–7 ans d'expérience pertinente.
  • Master: 1–3 ans d'expérience pertinente.

Responsibilities

  • Gère et améliore autonomement le cycle de vie des plateformes et services, du design au déploiement et à la retraite.
  • Conduit des revues de capacité et des post-mortems pour prévenir les incidents.
  • Conduit des solutions d'infrastructure à grande échelle et guide les équipes sur les meilleures pratiques DevOps.

Skills

Kubernetes
Containers
Python
Bash
Go
CI/CD
Cloud AWS
Observabilité

Education

Baccalauréat en informatique/sciences connues ou équivalent
Master préféré

Tools

CKA
Rancher
Portworx
Ansible
Terraform

Job description

Chez Roche, vous pouvez être vous-même et être apprécié pour les qualités uniques que vous apportez. Notre culture encourage l'expression personnelle, le dialogue ouvert et les connexions authentiques, où vous êtes valorisé, accepté et respecté pour ce que vous êtes, vous permettant de prospérer tant personnellement que professionnellement. Voici comment nous visons à prévenir, arrêter et guérir les maladies et à garantir à chacun l'accès aux soins de santé aujourd'hui et pour les générations à venir. Rejoignez Roche, où chaque voix compte.

La position

Kubernetes Reliability Engineer

A healthier future. It’s what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. That’s what makes us Roche.

The CaaS IT Infrastructure Engineer is a highly skilled expert responsible for solving complex business problems using advanced cloud native technologies. The engineer will build and maintain a Kubernetes-based infrastructure, enabling the modernization of business applications and processes.

This role combines software and systems engineering to optimize systems, increase efficiency, and eliminate operational work through automation.

The Opportunity

You will be part of the global CaaS infrastructure team at a leading healthcare company, working with members across different regions. The team's mandate is to deliver, maintain, and continuously improve a highly available Kubernetes platform across hybrid cloud deployments, including on-premise data centers and public clouds like AWS. In this role, you will apply software engineering principles to operations to build and run massively distributed, fault-tolerant systems, focusing heavily on automation, security, and observability.

  • Scope: Engages in and improves autonomously the whole lifecycle of platforms and services—from inception and design through deployment, operation, and retirement. Applies software engineering principles to build and manage large-scale IT infrastructure products, abstracting away complexity by providing self-service tools and APIs for developers. Designs, implements, and maintains CI/CD pipelines and develops self-healing features.
  • Service Reliability and Optimization: Focus on capacity planning and launch reviews for services before they go live. Perform blameless postmortems and proactive identification of potential outages to foster iterative improvements
  • Accountability/Problem Solving: Resolves complex problems in a global Kubernetes-based infrastructure through in-depth evaluation of variable factors, including inter-organizational impact, balanced with effective consultative engagement of key stakeholders. Leads end-to-end design of infrastructure solutions and maintains component standards. Evaluates promising solutions via Proof of Concept (PoCs) and feasibility studies across multiple areas, and serves as an internal escalation point for major incidents
  • Stakeholder Management: Acts as a bridge between engineering and operations. Communicates and presents complex information and potential solutions to cross-functional teams and the business in non-technical terms. Represents the organization as a prime contact on initiatives and interacts with senior internal and external personnel. Uses deep knowledge to influence IT infrastructure vendor product evaluations and collaborates with multiple IT partners (e.g. Enterprise Architects, Solution Owners) to integrate feedback. Mentors and shares DevOps culture, guiding developers on how to create and deploy cloud-native applications
  • Impact/Strategy: Provides technical leadership and direction for small-to-medium sized initiatives (projects, lifecycle work, PoCs). Ensures solutions comply with Quality/Regulatory standards and that designs adhere to the organization’s Technical Architecture Framework (TAF) policies and directions. Assists in planning technology projects, estimating engineering resources, dependencies, risks and timelines for successful delivery
  • Complexity: Demonstrates the ability to lead geo-distributed initiatives across different locations and cultural backgrounds through influence and mentorship, providing specialized guidance to drive success and cohesion.
  • Business/Technical ability: Applies extensive cloud native technical expertise, acting as a recognized expert in Kubernetes and maintaining in-depth knowledge across related cloud native technologies (containers, AWS, etc.). Demonstrates a detailed understanding of how IT infrastructure impacts respective Roche business processes and outcomes.
Who you are :
Education / Experience
  • Bachelor’s degree in Computer Science, Mathematics, Physics or related field, and 2-5 years of relevant experience.
  • Without degree: 4-7 years of relevant experience.
  • Master’s degree: 1-3 years of relevant experience.
Technical Skills
  • Kubernetes & Containers: Strong hands-on experience navigating, managing, and hardening Kubernetes clusters and containers, including knowledge of distributed storage. A Certified Kubernetes Administrator (CKA) certification is a strong plus. Knowledge of tools like Rancher or Portworx is beneficial
  • Infrastructure as Code (IaC): Hands-on experience delivering and managing infrastructure automation using tools like Ansible and Terraform
  • Scripting & Software Engineering: Proficiency in scripting and programming languages, primarily Python, Bash, or Go, including experience with test automation (e.g., pytest) and APIs deployment and management
  • CI/CD Tools: Expert knowledge of implementing software delivery pipelines using tools (e.g., Jenkins, Rundeck, or GitLab)
  • Systems & Networking: Strong understanding of Linux operating systems and core networking principles, including DNS, load balancing, firewalls, routing, and service meshes.
  • Observability: Experience configuring logging, metrics, and monitoring tools, specifically focusing on setting up alerts based on symptoms rather than waiting for system outages
  • Cloud Infrastructure: Experience with public cloud platforms, with a strong preference for AWS, specifically involving managed services for compute, networking, security, and identity (e.g., EKS, VPC, IAM).
General and Operational Knowledge
  • You have a proven experience applying best practices in an always-up, always-available service environment utilizing Scrum and Agile methodologies
  • You demonstrate a deep understanding of Technical Architecture Frameworks (TAF) and Quality/Regulatory compliance standards
Additional Qualifications
  • You have excellent problem-solving skills, decision-making ability, and sound judgment
  • You demonstrated a strong team-oriented mindset with the ability to function independently with low supervision and navigate ambiguity.
  • You are highly fluent in oral and written English communication skills are required.
  • You have the ability to work across multiple time zones.
  • You demonstrate strong customer & delivery focus with the ability to act as an analyst, seamlessly transforming complex stakeholder needs into actionable technical requirements.
  • You possess strong practice of sustainable incident response, including managing ITSM processes and leading audit evidence collection.

Relocation benefits are not available for this position.

The expected salary range for this position based on the primary location of Mississauga is 105,560.00 and 138,547.50 of hiring range. Actual pay will be determined based on experience, qualifications, and other job-related factors as determined by the company.

Nous utilisons l’intelligence artificielle pour filtrer, évaluer ou sélectionner les candidatures pour ce poste.

Cet affichage est pour un poste vacant existant chez Hoffmann-La Roche Ltée.

Qui nous sommes

Un avenir plus sain nous pousse à innover. Ensemble, plus de 100 000 employés à travers le monde sont dédiés à faire progresser la science et à garantir à chacun l'accès aux soins de santé aujourd'hui et pour les générations à venir. Nos efforts aboutissent à plus de 26 millions de personnes traitées avec nos médicaments et plus de 30 milliards de tests réalisés avec nos produits de Diagnostique. Nous nous encourageons mutuellement à explorer de nouvelles possibilités, à favoriser la créativité et à conserver nos grandes ambitions, afin de fournir des solutions de santé qui changent des vies et ont un impact mondial.

Construisons ensemble un avenir plus sain.

Roche est un employeur offrant l'équité en matière d'emploi.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Principal DSX Data Scientist
Senior Principal DSX Data Scientist

Roche Holding AG • Mississauga

On-site
CAD 149,000 - 195,000
IT Workplace Professional
IT Workplace Professional

Roche Holding AG • Laval (administrative region)

On-site
CAD 66,000 - 87,000
Healthcare System Partner (HSP) - West
Healthcare System Partner (HSP) - West

Roche Holding AG • Mississauga

On-site
CAD 137,000 - 180,000
IT Workplace Professional
IT Workplace Professional

Roche • Laval (administrative region)

On-site
CAD 66,000 - 87,000
Procurement Manager - Corporate Services
Procurement Manager - Corporate Services

Roche Holding AG • Mississauga

On-site
CAD 115,000 - 151,000
IT Workplace Professional
IT Workplace Professional

F. Hoffmann-La Roche AG • Laval (administrative region)

On-site
CAD 66,000 - 87,000
2027 Winter Internship - External Quality System Intern
2027 Winter Internship - External Quality System Intern

Roche Holding AG • Mississauga

On-site
CAD 57,000 - 75,000
Global Clinical Operations Excellence Leader
Global Clinical Operations Excellence Leader

Roche Holding AG • Mississauga

On-site
CAD 137,000 - 180,000
PT Learning Business Partnering Lead
PT Learning Business Partnering Lead

Roche Holding AG • Mississauga

On-site
CAD 160,000 - 210,000
Kubernetes Reliability Engineer
Kubernetes Reliability Engineer

F. Hoffmann-La Roche AG • Mississauga

On-site
CAD 106,000 - 139,000