Orchestration Workload Engineer - ACE - AI Factory

Roche

Kaiseraugst

Vor Ort

CHF 150.000 - 190.000

Vollzeit

Vor 10 Tagen
Bewerbungsgenerator

Eine maßgeschneiderte Bewerbung für diese Stelle — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Roche Kaiseraugst seeks a Workload Orchestration Engineer to own and advance SLURM-based scheduling across HPC and AI environments. You will drive resource optimization for multi-node CPU/GPU, ensuring high availability and efficient execution of research workloads.

You will lead cross‑functional initiatives, mentor engineers, and collaborate with Observability to deploy telemetry dashboards that monitor queue times and hardware utilization.

Qualifikationen

  • Bachelor's or advanced degree in Computer Science, Applied Mathematics, or Computational Engineering.
  • Extensive experience in workload scheduling and SLURM administration.
  • Proven track record leading complex technical initiatives.
  • Experience in life sciences, pharmaceutical R&D, or high‑performance computing environments.

Aufgaben

  • Lead SLURM architecture, tuning, and policy design.
  • Bridge HPC and cloud-native ecosystems with Kubernetes integrations.
  • Provide technical leadership and mentorship across cross-functional teams.
  • Partner with Observability to implement telemetry dashboards for efficiency.

Kenntnisse

SLURM administration
Workload scheduling
Multi-tenant cluster optimization
Technical leadership & mentoring
Life sciences / pharmaceutical R&D

Ausbildung

Bachelor's or higher in Computer Science / Applied Mathematics / Computational Engineering

Tools

SLURM
Kubernetes
Singularity/Apptainer

Jobbeschreibung

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

The Position

As a Workload Orchestration Engineer within the Accelerated Compute Engineering (ACE) team, you will be recognised internally as an expert in workload orchestration, owning and advancing our scheduler tech stack across our High-Performance Computing (HPC) platforms. With the rapid expansion of our compute infrastructure, your broad expertise will drive the efficient scheduling, policy management, and resource optimization of our multi-node CPU and GPU environments.

In this role, you will use your expertise to bridge traditional scientific computing with modern AI paradigms, while acting as a coach and mentor to help colleagues develop technical expertise. You will solve unique, unprecedented scheduling and infrastructure challenges that directly impact Roche's compute architecture, ensuring our researchers, data scientists, and engineers can execute compute workloads reliably, efficiently, and successfully.

Hosting and Infrastructure (HI) provides mission-critical on-premises infrastructure, cloud hosting, connectivity, and technology products that enable all functions at every Roche site to develop, innovate, connect, and deliver compliant digital products across the Roche Enterprise.

The Value Streams - Accelerated Compute Engineering (ACE) Team acts as a center of excellence and delivery for High Performance Compute and AI Infrastructure across Roche. This team facilitates seamless onboarding and adoption for business vertical customers needing accelerated compute-helping infrastructure consumers optimize for high availability, seamless data transfer, flexibility, speed, and the rapidly changing needs of AI to achieve rapid time-to-value.

The Opportunity
  • SLURM Architecture & Ecosystem Leadership

    Serve as the internal expert on the SLURM Workload Manager, architecting, scaling, and maintaining SLURM across heterogeneous HPC (and AI environments) to ensure high availability and dynamic resource distribution.

    Design and tune advanced SLURM configurations, including custom plugin integration, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and complex QoS/fair-share policies.

    Bridge HPC and cloud-native ecosystems by evaluating and implementing integrations between SLURM, Kubernetes, and orchestration platforms (e.g., SLURM Slinky or Run:ai) to streamline job submission workflows across architectures.

  • Hybrid Workload & Kubernetes Integration

    Integrate containerization standards across SLURM (using Singularity/Apptainer) while maintaining operational familiarity with Kubernetes container orchestration to support hybrid AI/HPC workloads.

    Solve unique, unprecedented multi-tenant bottlenecks, such as GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures.

  • Technical Leadership, Mentorship & Governance

    Lead large, global cross-functional initiatives across ACE, infrastructure, platform, scientific computing, and AI teams to establish workload orchestration standards, policies, and architectural patterns across Roche compute environments.

    Act as a technical mentor and coach for junior and mid-level engineers, driving skill development and continuous learning across the chapter.

    Partner with Observability Engineers to establish deep telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization, utilizing configuration-as-code to deploy policies uniformly.

Who You Are
  • Bachelor's or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline.
  • Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization.
  • Demonstrated track record of leading complex technical initiatives and mentoring engineering peers.
  • Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments.
SLURM Architecture & Optimization

Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments, including scheduler/backfill tuning, partition and topology design, priority/fair-share/QoS policies, cgroups, HA architecture, plugin integration, and GRES/TRES modeling for GPUs and specialized resources.

SLURM Operations, Accounting & Observability

Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, util

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Orchestration Workload Engineer - ACE - AI Factory
Orchestration Workload Engineer - ACE - AI Factory

Genentech • Kaiseraugst

Vor Ort
CHF 120.000 - 160.000
Orchestration Workload Engineer - ACE - AI Factory
Orchestration Workload Engineer - ACE - AI Factory

F. Hoffmann-La Roche AG • Kaiseraugst

Vor Ort
CHF 120.000 - 180.000
SLURM & AI HPC Orchestration Engineer
SLURM & AI HPC Orchestration Engineer

F. Hoffmann-La Roche AG • Kaiseraugst

Vor Ort
CHF 120.000 - 180.000
SLURM & HPC Orchestration Engineer for AI Workloads
SLURM & HPC Orchestration Engineer for AI Workloads

Roche • Kaiseraugst

Vor Ort
CHF 150.000 - 190.000
AI/HPC Workload Orchestration Engineer
AI/HPC Workload Orchestration Engineer

Genentech • Kaiseraugst

Vor Ort
CHF 120.000 - 160.000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA Gruppe • Zürich

Vor Ort
CHF 120.000 - 180.000
HPC Systems Engineer - SLURM & Large-Scale Compute
HPC Systems Engineer - SLURM & Large-Scale Compute

Big Science Sweden • Genf

Vor Ort
CHF 70.000 - 100.000
High-Performance Compute (HPC) Engineer
High-Performance Compute (HPC) Engineer

Big Science Sweden • Genf

Vor Ort
CHF 70.000 - 100.000
Senior Technical Specialist
Senior Technical Specialist

Elan Personal AG • Basel

Vor Ort
CHF 120.000 - 190.000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Schweiz

Vor Ort
CHF 150.000 - 210.000