Workload Orchestration Engineer

F. Hoffmann-La Roche AG

Madrid

Presencial

EUR 70.000 - 110.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

F. Hoffmann-La Roche AG in Madrid seeks a Workload Orchestration Engineer within the ACE team to oversee and advance our workload orchestration stack for HPC and AI Factory platforms.

You will deploy, configure, and tune SLURM across clusters and run Run:ai to enable fractional GPU allocation and dynamic scheduling for multi-tenant workloads. You will define containerized execution using Singularity/Apptainer and Enroot, optimize QoS and queues, and collaborate with Observability to build

Formación

  • 5+ years of systems engineering experience focusing on workload scheduling and multi-tenant cluster optimization.
  • Deep familiarity with Enterprise Linux and distributed systems architectures.
  • Expert-level proficiency in SLURM administration, partitions, accounting and plugins.
  • Hands-on experience with Run:ai, Kubernetes, and GPU scheduling paradigms.
  • Automation of scheduler configurations and telemetry collection.

Responsabilidades

  • Orchestration Stack Deployment & Governance: design, deploy, and maintain SLURM across HPC cluster architectures.
  • Deploy and manage Run:ai as the core AI orchestration layer for fractional GPU allocation.
  • Integrate SLURM Slinky with Kubernetes where needed to bridge HPC and container-based workflows.
  • Containerization & Workload Optimization: define best practices with Singularity/Apptainer/Enroot.
  • Profile and tune queues, QoS, and fair-share policies for multi-tenant efficiency.
  • Platform Reliability & Telemetry: implement monitoring, dashboards, and telemetry for scheduler performance.
  • Troubleshoot distributed training, MPI/NCCL issues and driver compatibility.

Conocimientos

Workload scheduling
Resource management
Cluster optimization
Linux administration
Automation
Cross-team collaboration

Educación

Bachelor’s or advanced degree in Computer Science or related field

Herramientas

SLURM
Run:ai
Kubernetes
Singularity/Apptainer
Enroot

Descripción del empleo

The Position

As a Workload Orchestration Engineer within the Accelerated Compute Engineering (ACE) team, you will be responsible for overseeing and advancing our workload orchestration tech stack across both our High-Performance Computing (HPC) and industry-leading AI Factory platforms. With the rapid expansion of our compute infrastructure, efficiently scheduling, managing, and maximizing the utilization of our CPU and GPU environments is paramount. You will own the deployment, configuration, and fine-tuning of orchestration platforms that schedule massive, parallel computational workloads. By implementing robust scheduling policies for traditional scientific workflows and modern containerized AI workloads, you will bridge the gap between heavy compute capacity and efficient execution. Your work will directly ensure that Roche’s researchers, data scientists, and engineers can seamlessly run large-scale AI model training and computational science simulations at scale.

Description of the area

Hosting and Infrastructure (HI) provides mission-critical on-premise infrastructure, cloud hosting, connectivity, and technology products that enable all functions at every Roche site to develop, innovate, connect, and deliver compliant digital products across the Roche Enterprise.

The Value Streams

- Accelerated Compute Engineering (ACE) Team is focused on driving both customer success and platform success by acting as a center of excellence and delivery for the High Performance Compute and AI Infrastructure supporting AI and HPC use cases across Roche. This team facilitates seamless onboarding and adoption for business vertical customers needing accelerated compute—helping those infrastructure consumers with needs optimized for high availability, seamless data transfer, flexibility, speed, and the rapidly changing needs of AI—helping achieve rapid time-to-value.

Job Responsibilities

Orchestration Stack Deployment & Governance Design, implement, and maintain the SLURM Workload Manager ecosystem across our HPC cluster architectures, ensuring high availability and optimal resource distribution. Deploy and manage Run:ai as the core orchestration and virtualization layer for the AI Factory, enabling fractional GPU allocation and dynamic resource allocation. Evaluate, architect, and implement SLURM Slinky integrations where required to seamlessly bridge Kubernetes-based AI orchestration with traditional HPC cluster resources. Containerization & Workload Optimization Define best practices and frameworks for containerized scientific execution, utilizing Singularity/Apptainer and/or Enroot to provide secure, reproducible performance environments for HPC. Translate user and workload requirements into optimized scheduling parameters (e.g., topology-aware scheduling, multi-node scaling). Actively profile and tune scheduling queues, quality-of-service (QoS) parameters, and fair-share policies to maximize multi-tenant efficiency. Platform Reliability & Telemetry Partner with Observability Engineers to implement continuous monitoring, telemetry, and reporting dashboards to track scheduler efficiency, queue wait times, and hardware utilization rates. Troubleshoot complex workload failures, including distributed training synchronization issues, MPI communication bottlenecks, and driver incompatibilities. Maintain configuration-as-code models for the scheduling tier, leveraging automation to deploy cluster policies uniformly.

Qualifications

Education / Experience Bachelor’s or an advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a similar technical discipline. 5+ years of systems engineering experience, with a heavy emphasis on workload scheduling, resource management, and cluster optimization for multi-tenant environments. Deep technical familiarity with Enterprise Linux operating systems and distributed systems architecture. HPC Scheduling & Tooling: Expert-level proficiency in administering SLURM, including complex partition designs, accounting, and plug-in management. Highly proficient with Singularity for container runtime execution. AI Orchestration: Hands-on experience or deep architectural understanding of Run:ai, Kubernetes, and containerized GPU scheduling paradigms. Infrastructure Literacy: Solid understanding of high-speed interconnects (InfiniBand, RoCE) and multi-node communication architectures (MPI, NCCL) as they relate to job placement. Automation: Proficiency in automating scheduler configurations and telemetry gathering, or infrastructure automation tooling. Leadership & Mindset: Lean & Agile Mindset: Highly focused on driving efficiency, reducing idle compute time, and creating frictionless pathways for user workload submissions. Collaboration & Advocacy: Outstanding capability to translate scientific and AI model workflow challenges into scalable scheduler configurations. Intellectual Curiosity: A strong passion for remaining ahead of industry trends regarding GPU slicing, fractionalization, and the convergence of AI workloads with traditional HPC schedulers.

Who we are

A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact. Let’s build a healthier future, together.

Roche is an Equal Opportunity Employer. We believe it’s urgent to deliver medical solutions right now – even as we develop innovations for the future. We are passionate about transforming patients’ lives. We are courageous in both decision and action. And we believe that good business means a better world. That is why we come to work each day. We commit ourselves to scientific rigor, unassailable ethics, and access to medical innovations for all. We do this today to build a better tomorrow. We are proud of who we are, what we do, and how we do it. We are many, working as one across functions, across companies, and across the world. We are Roche.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Finops Engineer - Hosting and Infrastructure
Finops Engineer - Hosting and Infrastructure

F. Hoffmann-La Roche AG • Madrid

Presencial
EUR 90.000 - 140.000
DevOps Infrastructure Engineer
DevOps Infrastructure Engineer

F. Hoffmann-La Roche AG • Madrid

Presencial
EUR 60.000 - 90.000
Finops Engineer - Hosting and Infrastructure
Finops Engineer - Hosting and Infrastructure

Roche • Madrid

Presencial
EUR 90.000 - 120.000
Cloud Engineer - Data Platforms
Cloud Engineer - Data Platforms

F. Hoffmann-La Roche AG • Madrid

Presencial
EUR 65.000 - 95.000
Workflow & IT Technical Specialist
Workflow & IT Technical Specialist

F. Hoffmann-La Roche AG • Sant Cugat del Vallès

Presencial
EUR 45.000 - 65.000
Cloud DevOps Infrastructure Engineer - Roche Cloud Platform GCP (Evening Shift)
Cloud DevOps Infrastructure Engineer - Roche Cloud Platform GCP (Evening Shift)

Roche • Madrid

Presencial
EUR 85.000 - 125.000
Senior HPC & AI Workload Orchestration Engineer
Senior HPC & AI Workload Orchestration Engineer

F. Hoffmann-La Roche AG • Madrid

Presencial
EUR 70.000 - 110.000
Team Lead Engineering - Content Search and Knowledge Management 2
Team Lead Engineering - Content Search and Knowledge Management 2

F. Hoffmann-La Roche AG • Madrid

Presencial
EUR 90.000 - 130.000
IT Software Engineer - RDT Engineering Excellence & Experience
IT Software Engineer - RDT Engineering Excellence & Experience

Roche • Madrid

Presencial
EUR 70.000 - 100.000
Senior Project Manager - RDT Data
Senior Project Manager - RDT Data

F. Hoffmann-La Roche AG • Madrid

Presencial
EUR 90.000 - 130.000