Orchestration Workload Engineer - ACE - AI Factory

Roche Holding AG

Kaiseraugst

Vor Ort

CHF 180.000 - 250.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Eine vollständige Bewerbung in einer Minute — Lebenslauf und Anschreiben, maßgeschneidert und versandbereit.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Roche Holding AG sucht eine erfahrene Fachkraft als Workload Orchestration Engineer, um die SLURM-Architektur in HPC- und KI-Umgebungen zu führen. Sie arbeiten an der Optimierung von Scheduling, QoS und Multi-Tenant-Betrieb.

Zudem bridgen Sie HPC mit Kubernetes für hybride AI/HPC-Workflows und übernehmen fachliche Führungsverantwortung. Sie gestalten Governance, Mentoring und Telemetrie-Dashboards und arbeiten eng mit globalen Teams zusammen, um Leistungsfähigkeit, Verfügbarkeit und

Qualifikationen

  • Extensive systems engineering experience with workload scheduling and SLURM.
  • Proven track record leading technical initiatives and mentoring peers.
  • Experience in life sciences, pharmaceutical R&D, or high-performance research.
  • Strong troubleshooting of complex HPC/AI workloads.
  • Knowledge of cloud-native orchestration and containerized workloads.

Aufgaben

  • Own and advance SLURM architecture across HPC and AI environments.
  • Design tuning of SLURM configurations and QoS policies.
  • Bridge HPC with Kubernetes for hybrid AI/HPC workflows.
  • Provide technical leadership and mentoring across global teams.
  • Develop telemetry dashboards with Observability Engineers.

Kenntnisse

Workload scheduling
SLURM administration
Mentoring
Cross-functional collaboration
Problem solving

Ausbildung

Bachelor’s degree in Computer Science or related field
Advanced degree in a related field

Tools

Kubernetes
Singularity/Apptainer
Docker
Ansible
Terraform
InfiniBand/RoCE networking
MPI/NCCL

Jobbeschreibung

Bei Roche kannst du ganz du selbst sein und wirst für deine einzigartigen Qualitäten geschätzt. Unsere Kultur fördert persönlichen Ausdruck, offenen Dialog und echte Verbindungen. Hier wirst du für das, was du bist, wertgeschätzt, akzeptiert und respektiert. Dies schafft ein Umfeld, in dem du sowohl persönlich als auch beruflich wachsen kannst. Gemeinsam wollen wir Krankheiten vorbeugen, stoppen und heilen und sicherstellen, dass jeder Zugang zur Gesundheitsversorgung hat – heute und in Zukunft. Werde Teil von Roche, wo jede Stimme zählt.

Die Position

As a Workload Orchestration Engineer within the Accelerated Compute Engineering (ACE) team, you will be recognised internally as an expert in workload orchestration, owning and advancing our scheduler tech stack across our High-Performance Computing (HPC) platforms. With the rapid expansion of our compute infrastructure, your broad expertise will drive the efficient scheduling, policy management, and resource optimization of our multi-node CPU and GPU environments.

In this role, you will use your expertise to bridge traditional scientific computing with modern AI paradigms, while acting as a coach and mentor to help colleagues develop technical expertise. You will solve unique, unprecedented scheduling and infrastructure challenges that directly impact Roche’s compute architecture, ensuring our researchers, data scientists, and engineers can execute compute workloads reliably, efficiently, and successfully.

Hosting and Infrastructure (HI) provides mission-critical on-premises infrastructure, cloud hosting, connectivity, and technology products that enable all functions at every Roche site to develop, innovate, connect, and deliver compliant digital products across the Roche Enterprise.

The Value Streams - Accelerated Compute Engineering (ACE) Team acts as a center of excellence and delivery for High Performance Compute and AI Infrastructure across Roche. This team facilitates seamless onboarding and adoption for business vertical customers needing accelerated compute—helping infrastructure consumers optimize for high availability, seamless data transfer, flexibility, speed, and the rapidly changing needs of AI to achieve rapid time-to-value.

The Opportunity:

SLURM Architecture & Ecosystem Leadership

  • Serve as the internal expert on the SLURM Workload Manager, architecting, scaling, and maintaining SLURM across heterogeneous HPC (and AI environments) to ensure high availability and dynamic resource distribution.

  • Design and tune advanced SLURM configurations, including custom plugin integration, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and complex QoS/fair-share policies.

  • Bridge HPC and cloud-native ecosystems by evaluating and implementing integrations between SLURM, Kubernetes, and orchestration platforms (e.g., SLURM Slinky or Run:ai) to streamline job submission workflows across architectures.

Hybrid Workload & Kubernetes Integration

  • Integrate containerization standards across SLURM (using Singularity/Apptainer) while maintaining operational familiarity with Kubernetes container orchestration to support hybrid AI/HPC workloads.

  • Solve unique, unprecedented multi-tenant bottlenecks, such as GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures.

Technical Leadership, Mentorship & Governance

  • Lead large, global cross-functional initiatives across ACE, infrastructure, platform, scientific computing, and AI teams to establish workload orchestration standards, policies, and architectural patterns across Roche compute environments.

  • Act as a technical mentor and coach for junior and mid-level engineers, driving skill development and continuous learning across the chapter.

  • Partner with Observability Engineers to establish deep telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization, utilizing configuration-as-code to deploy policies uniformly.

Who You Are:
  • Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline.

  • Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization.

  • Demonstrated track record of leading complex technical initiatives and mentoring engineering peers.

  • Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments.

  • SLURM Architecture & Optimization: Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments, including scheduler/backfill tuning, partition and topology design, priority/fair-share/QoS policies, cgroups, HA architecture, plugin integration, and GRES/TRES modeling for GPUs and specialized resources.

  • SLURM Operations, Accounting & Observability: Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, utilization analysis, and troubleshooting complex controller, database, node, and workload interactions.

  • Kubernetes & Container Knowledge: Hands-on experience with Kubernetes fundamentals and container runtimes (Singularity, Apptainer, Enroot, Docker) within an HPC context.

  • AI Infrastructure & Interconnects: Deep familiarity with GPU scheduling (NVIDIA MIG, fractionalization), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL).

  • Automation: Advanced proficiency with Infrastructure-as-Code (Ansible, Terraform) to automate scheduler deployments, configuration drift management, and telemetry pipelines.

  • Broad Platform Expertise: Apply broad knowledge across HPC, AI infrastructure, Kubernetes, containers, networking/interconnects, observability, automation, and capacity management to solve orchestration problems spanning multiple technology domains.

  • Domain Expertise & Problem Solving: Proven ability to troubleshoot complex, unprecedented failure modes at the intersection of hardware, OS, schedulers, and workloads.

  • Coaching & Collaboration: Strong leadership presence with a dedication to mentoring colleagues, driving technical standards, and collaborating with global cross-functional teams.

  • Cross-Organizational Coordination: Collaborative team player with demonstrated ability to coordinate initiatives across diverse global business units, IT functions, and scientific research stakeholders.

  • Strategic Vision: Passion for guiding the convergence of traditional HPC schedulers like SLURM with cloud-native, Kubernetes-driven AI workflows.

#RDT2026

Where pay transparency applies, details are provided based on the primary posting location. For this role, the primary location is Kaiseraugst. If you are interested in additional locations where the role may be available, we will provide the relevant compensation details later in the hiring process.

Wer wir sind

Eine gesündere Zukunft treibt uns zur Innovation an. Mehr als 100.000 Mitarbeiter weltweit arbeiten gemeinsam daran, wissenschaftliche Fortschritte zu erzielen und sicherzustellen, dass jeder Zugang zur Gesundheitsversorgung hat – heute und für zukünftige Generationen. Durch unser Engagement werden über 26 Millionen Menschen mit unseren Medikamenten behandelt und mehr als 30 Milliarden Tests mit unseren Diagnostik-Produkten durchgeführt. Wir ermutigen uns gegenseitig, neue Möglichkeiten zu erkunden, Kreativität zu fördern und hohe Ziele zu setzen, um lebensverändernde Gesundheitslösungen zu liefern.

Gemeinsam können wir eine gesündere Zukunft gestalten.

Roche ist ein Arbeitgeber, der die Chancengleichheit fördert.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Team Lead - Workflow Innovation & Management
Team Lead - Workflow Innovation & Management

Roche Holding AG • Zug

Vor Ort
CHF 150.000 - 210.000
Contract Manufacturing Specialist
Contract Manufacturing Specialist

Roche Holding AG • Kaiseraugst

Vor Ort
CHF 150.000 - 190.000
Medical Device Technical Lead
Medical Device Technical Lead

Roche Holding AG • Basel

Vor Ort
CHF 240.000 - 320.000
Senior DMPK-PD Project Leader
Senior DMPK-PD Project Leader

Roche Holding AG • Basel

Vor Ort
CHF 200.000 - 320.000
Global Diagnostics Partnering Strategy & Transactions Lead
Global Diagnostics Partnering Strategy & Transactions Lead

Roche Holding AG • Zug

Vor Ort
CHF 180.000 - 240.000
Internship for students as Global Clinical Artwork Manager - RiKO (from February 2027, 12 months)
Internship for students as Global Clinical Artwork Manager - RiKO (from February 2027, 12 months)

Roche Holding AG • Kaiseraugst

Vor Ort
CHF 13.000 - 18.000
Operations Portfolio Lead BGE - Blood Gas & Electrolyte Portfolio
Operations Portfolio Lead BGE - Blood Gas & Electrolyte Portfolio

Roche Holding AG • Zug

Vor Ort
CHF 180.000 - 240.000
DMPK-PD Project Leader
DMPK-PD Project Leader

Roche Holding AG • Basel

Vor Ort
CHF 180.000 - 240.000
Lehrstelle als Informatiker:in EFZ Applikationsentwicklung (ab August 2027)
Lehrstelle als Informatiker:in EFZ Applikationsentwicklung (ab August 2027)

Roche Holding AG • Basel

Vor Ort
CHF 13.000 - 17.000
Internationale Umgebung
Praxisorientierte Werkschule
Moderne Anlagen und Ausstattung
+1
Technischer Experte für Reinstmedien Anlagen
Technischer Experte für Reinstmedien Anlagen

F. Hoffmann-La Roche AG • Kaiseraugst

Vor Ort
CHF 90.000 - 120.000