Senior Site Reliability Engineer

F. Hoffmann-La Roche AG

Sant Cugat del Vallès

Presencial

EUR 50.000 - 75.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

F. Hoffmann-La Roche AG is seeking a Site Reliability Engineer in Sant Cugat del Vallès. The candidate will design, build, and scale distributed systems for healthcare innovation. Responsibilities include defining SLIs/SLOs, conducting reliability reviews, and participating in a structured on-call rotation.

Required qualifications include a Bachelor’s degree in a related field and experience in cloud management with AWS or Azure. The company offers a collaborative work environment focused on operational excellence and reliability improvements.

Formación

  • Bachelor’s degree or equivalent professional experience is required.
  • Experience in site reliability engineering or software engineering with production on-call experience is essential.
  • Proficiency in scripting languages for automation (Python, etc.) is necessary.
  • Excellent communication and teamwork skills are required.

Responsabilidades

  • Define and implement SLIs, SLOs, and error budgets.
  • Conduct reliability reviews and design fault-tolerant architectures.
  • Participate in 24/7 on-call rotation and root-cause analysis.
  • Automate operational processes and improve CI/CD reliability.

Conocimientos

Site reliability engineering
AWS
Azure
Python
Cloud resources management
Observability tools
Incident management tools
Kubernetes

Educación

Bachelor’s degree in computer science, engineering, or a related field

Herramientas

Terraform

Descripción del empleo

Position Overview

We are building a global Site Reliability Engineering (SRE) team to support critical commercial and internal platforms and applications. As an SRE, you will design, build, and scale reliable distributed systems that power healthcare innovation worldwide. The role focuses on reliability, scalability, automation, and operational excellence and includes participation in a structured on‑call rotation.

Core Responsibilities
  • Define and implement SLIs, SLOs, and error budgets with product and engineering teams.
  • Conduct reliability reviews for new and existing services.
  • Design scalable, fault‑tolerant architectures in AWS and Azure environments.
  • Lead capacity planning, performance and cost optimization initiatives.
  • Improve system resilience through automation and self‑healing patterns.
  • Drive organizational observability maturity (metrics, logs, traces, alert quality).
  • Perform complex root‑cause analysis and drive rapid mitigation.
  • Participate in blameless post‑mortems and follow‑through.
  • Improve MTTR, reduce incident frequency, and elevate production standards.
  • Collaborate seamlessly with engineering teams to enable timely and effective resolutions.
  • Handle requests and incidents, create and maintain runbooks.
  • Participate in a structured 24/7 on‑call rotation.
  • Reduce operational toil through tooling and automation (Python or similar).
  • Improve CI/CD reliability and deployment safety mechanisms.
  • Build and maintain infrastructure‑as‑code (Terraform or equivalent).
  • Enhance Kubernetes platform reliability (EKS, AKS, or similar).
  • Partner with business, engineering, security, and cloud teams to embed reliability early in the software development life cycle.
  • Mentor mid‑level engineers and help shape SRE best practices.
  • Champion a culture of ownership, accountability, and continuous improvement.
Qualifications
  • Bachelor’s degree in computer science, engineering, or a related field, or equivalent professional experience.
  • Experience in site reliability engineering, software engineering, or related fields with production on‑call experience.
  • Solid experience with AWS and/or Azure, including setting up, monitoring, and maintaining cloud resources (incl. Kubernetes, EKS, AKS, GKE).
  • Proficiency with observability tools.
  • Hands‑on experience with incident management tools.
  • Proficiency in scripting languages for automation purposes (Python, etc.).
  • Demonstrated proficiency in troubleshooting, especially in cloud and distributed system environments.
  • Excellent communication, teamwork and documentation skills, with a proactive and self‑motivated approach to improving system reliability and operational efficiencies.
  • Proficient spoken and written English communication.
Location

Primary location: Sant Cugat del Vallès. Additional locations may be available.

Equal Opportunity Employer

Roche is an Equal Opportunity Employer. We believe it’s urgent to deliver medical solutions right now – even as we develop innovations for the future. We are passionate about transforming patients’ lives. We are courageous in both decision and action. And we believe that good business means a better world. We are committed to scientific rigor, unassailable ethics, and access to medical innovations for all. We are proud of who we are, what we do, and how we do it. We are many, working as one across functions, across companies, and across the world.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior SRE: Global Reliability & Cloud Automation
Senior SRE: Global Reliability & Cloud Automation

F. Hoffmann-La Roche AG • Sant Cugat del Vallès

Presencial
EUR 50.000 - 75.000
Senior DevOps Automation Engineer
Senior DevOps Automation Engineer

Roche • Sant Cugat del Vallès

Híbrido
EUR 50.000 - 80.000
Brand new Apple hardware
Fitness benefits
Public transport benefits
+3
OSS Lead - RDT Quality, Risk & Compliance
OSS Lead - RDT Quality, Risk & Compliance

F. Hoffmann-La Roche AG • Madrid

Presencial
EUR 90.000 - 120.000
Facilities Engineer Sant Cugat Site
Facilities Engineer Sant Cugat Site

F. Hoffmann-La Roche AG • Sant Cugat del Vallès

Presencial
EUR 55.000 - 85.000
Senior Site Reliability Engineer - Logging & Monitoring
Senior Site Reliability Engineer - Logging & Monitoring

Swiss Re • Madrid

Presencial
EUR 60.000 - 100.000
OSS Lead - RDT Quality, Risk & Compliance
OSS Lead - RDT Quality, Risk & Compliance

Roche • Madrid

Presencial
EUR 85.000 - 120.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Ribarroja del Turia

Híbrido
EUR 40.000 - 70.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps
Competitive compensation: USD-based pay with education, fitness, and team activity budgets
Exciting projects: Modern solutions with Fortune 500 and top product companies
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Madrid

Híbrido
EUR 45.000 - 60.000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Camlin Group • Málaga

Presencial
EUR 85.000 - 120.000
Senior DevOps Mobile Engineer
Senior DevOps Mobile Engineer

Roche • Sant Cugat del Vallès

Híbrido
EUR 50.000 - 70.000
Brand new Apple hardware
Fitness benefits
Public transport subsidies
+3