Engineer F/M - Development and optimization of a co-designed runtime system for running high performance machine learning chains

HiPEAC

Bordeaux

Sur place

EUR 45 000 - 70 000

Plein temps

Il y a 7 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.

Passez les filtres ATS

Résumé du poste

Inria Bordeaux seeks an engineer to develop and optimize a runtime system for machine learning chains, focusing on performance and energy efficiency through close collaboration with hardware and software teams.

The role will advance a co-designed runtime that integrates with CAMELIA's software stack, validating software in CI, documenting the work, and ensuring interoperability across components and configurations.

Qualifications

  • Excellent programming in C and C++.
  • Linux environments with Git and CI.
  • Parallel, distributed and GPU programming on HPC.
  • Machine learning chains for training/inference.
  • Experience in validation and verification of code.

Responsabilités

  • Develop and optimize the runtime system.
  • Define, tune and validate scheduling policies.
  • Validate software in CI processes.
  • Prepare, validate and maintain software documentation and interoperability.

Connaissances

C/C++ programming
Linux development
Parallel and GPU programming
ML pipelines
Code validation & verification

Outils

Git
CI frameworks
Linux
HPC platforms

Description du poste

The CAMELIA program (Hardware and Software Components for Advanced AI Accelerators) focuses on developing scalable, modular AI acceleration components and their associated software environment. It aims to create specialized accelerators that target key functions for executing AI models, using a joint hardware/software/application co-design approach. Anticipating breakthroughs in both hardware and software, the program aims to achieve significant performance gains—particularly in energy efficiency—compared to current industry solutions based on GPUs or NPUs.

One of the critical components of this co-designed architecture is the runtime system, whose role is precisely to drive the link between the application and the hardware. It monitors the effective use of computing resources and dynamically adjust the mapping of work onto such resources. It acts such as to minimize idle times and load imbalance, while ensuring the overlap of computations with non-computational activities such as load balancing and inputs/outputs.

The specific challenge for CAMELIA’s runtime system is to achieve this optimization work on extremely fine-grained work items. To successfully achieve this challenge, a tight collaboration with the compiler toolchain is necessary, in order to discover at run-time only the strictly dynamical information bits, and nevertheless enable building optimization decisions on a rich and detailed semantic context.

Mission

A prototype runtime system adapted to Program CAMELIA’s needs has been defined. The mission proposed to the person recruited will be to develop and optimize this prototype under the direction of permanent project members, to achieve a high level of performance in the execution of machine learning chains on the hardware infrastructure built by CAMELIA, and in relationship with the compiler toolchain assembled within this same project.

The person recruited will be responsible for developing and optimizing this prototype into a software foundation able to support the whole CAMELIA software stack, and to make the best use of the hardware platform of the project in all the diversity of its configurations and for assorted practical use cases. This optimization axis will aim to obtain a high level of performance, in terms of execution speed, but also with an energy footprint kept under control. For that, this work will be conducted in tight cooperation with project partners in charge of the development of the hardware and software modules, under a logic of synergistic co-design.

Moreover, the person recruited will be in charge of implementing the automated validation of software developments as a continuous integration framework, write software documentation, and technically supporting the project members to enforce and maintain a good interoperability of the runtime system with the other software and hardware components of Program CAMELIA.

Activities
Main activities
  • Software development of the runtime system;
  • Definition, tuning and optimisation of the proposed scheduling policies;
  • Validation of software developments in a continuous integration process;
  • Preparation, validation and maintenance of software documentation and of software interoperability with the other components of the project;
  • Evaluation and monitoring of performance on a set of use cases defined by the project.
Additional activities
  • Participation to the preparation of the reporting elements and communication elements of the project;
  • Participation to collective actions of the project: meetings, seminars, tutorials, workshops;
  • Collaboration with a pluridisciplinary community, at the interface between hardware and software.
Skills
Technical skills and level required
  • Excellent programming level in C and C++ languages;
  • Mastering of software development on Linux environments with the Git version control system, and with continuous integration frameworks;
  • Mastering of parallel programming, distributed programming and GPU programming on HPC platforms;
  • Mastering of machine learning chains’ usage for training and inference;
  • Experience in source code validation and verification.
Languages
  • Excellent written and spoken English level.
Relational skills
  • Personal implication and sense of the mission;
  • Ability to integrate and communicate in a plural scientific community.

Inria Bordeaux is a prominent research center within the French National Institute for Research in Digital Science and Technology, dedicated to advancing knowledge in digital technology in Bordeaux, France.

Metadata

Topics: Accelerators, Artificial intelligence, Deep learning, Energy efficiency / Low-power computing, High-performance computing, LLMs, Machine learning, Multicore / Manycore, Networking / Distributed computing, Optimization, Parallel computing, Performance engineering, Performance portability, Resource management / Scheduling, Runtime performance

Summary

Inria Bordeaux seeks an engineer to develop and optimize a runtime system for machine learning chains, focusing on performance and energy efficiency through close collaboration with hardware and software teams.

Engineer F/M - Development and optimization of a co-designed runtime system for running high performance machine learning chains

Full-time

Inria Bordeaux seeks an engineer to develop and optimize a runtime system for machine learning chains, focusing on performance and energy efficiency through close collaboration with hardware and software teams.

Context

The CAMELIA program (Hardware and Software Components for Advanced AI Accelerators) focuses on developing scalable, modular AI acceleration components and their associated software environment. It aims to create specialized accelerators that target key functions for executing AI models, using a joint hardware/software/application co-design approach. Anticipating breakthroughs in both hardware and software, the program aims to achieve significant performance gains—particularly in energy efficiency—compared to current industry solutions based on GPUs or NPUs.

One of the critical components of this co-designed architecture is the runtime system, whose role is precisely to drive the link between the application and the hardware. It monitors the effective use of computing resources and dynamically adjust the mapping of work onto such resources. It acts such as to minimize idle times and load imbalance, while ensuring the overlap of computations with non-computational activities such as load balancing and inputs/outputs.

The specific challenge for CAMELIA’s runtime system is to achieve this optimization work on extremely fine-grained work items. To successfully achieve this challenge, a tight collaboration with the compiler toolchain is necessary, in order to discover at run-time only the strictly dynamical information bits, and nevertheless enable building optimization decisions on a rich and detailed semantic context.

Mission

A prototype runtime system adapted to Program CAMELIA’s needs has been defined. The mission proposed to the person recruited will be to develop and optimize this prototype under the direction of permanent project members, to achieve a high level of performance in the execution of machine learning chains on the hardware infrastructure built by CAMELIA, and in relationship with the compiler toolchain assembled within this same project.

The person recruited will be responsible for developing and optimizing this prototype into a software foundation able to support the whole CAMELIA software stack, and to make the best use of the hardware platform of the project in all the diversity of its configurations and for assorted practical use cases. This optimization axis will aim to obtain a high level of performance, in terms of execution speed, but also with an energy footprint kept under control. For that, this work will be conducted in tight cooperation with project partners in charge of the development of the hardware and software modules, under a logic of synergistic co-design.

Moreover, the person recruited will be in charge of implementing the automated validation of software developments as a continuous integration framework, write software documentation, and technically supporting the project members to enforce and maintain a good interoperability of the runtime system with the other software and hardware components of Program CAMELIA.

Activities
Main activities
  • Software development of the runtime system;
  • Definition, tuning and optimisation of the proposed scheduling policies;
  • Validation of software developments in a continuous integration process;
  • Preparation, validation and maintenance of software documentation and of software interoperability with the other components of the project;
  • Evaluation and monitoring of performance on a set of use cases defined by the project.
Additional activities
  • Participation to the preparation of the reporting elements and communication elements of the project;
  • Participation to collective actions of the project: meetings, seminars, tutorials, workshops;
  • Collaboration with a pluridisciplinary community, at the interface between hardware and software.
Skills
Technical skills and level required
  • Excellent programming level in C and C++ languages;
  • Mastering of software development on Linux environments with the Git version control system, and with continuous integration frameworks;
  • Mastering of parallel programming, distributed programming and GPU programming on HPC platforms;
  • Mastering of machine learning chains’ usage for training and inference;
  • Experience in source code validation and verification.
Languages
  • Excellent written and spoken English level.
Relational skills
  • Personal implication and sense of the mission;
  • Ability to integrate and communicate in a plural scientific community.

The HiPEAC project has received funding from the European Union's Horizon Europe research and innovation funding programme under grant agreement number 101296676. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union. Neither the European Union nor the granting authority can be held responsible for them.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Engineer F/M - Adaptation of scientific codes to a task-based runtime system for exploiting high performance calibration and imaging pipelines in interferometric astronomy
Engineer F/M - Adaptation of scientific codes to a task-based runtime system for exploiting high performance calibration and imaging pipelines in interferometric astronomy

HiPEAC • Bordeaux

Sur place
EUR 55 000 - 75 000
Post-Doctoral Researcher F/M Dynamic Parallelization of Sparse Codes for High-Performance Computing and Machine Learning
Post-Doctoral Researcher F/M Dynamic Parallelization of Sparse Codes for High-Performance Computing and Machine Learning

HiPEAC • Lyon

Sur place
EUR 42 000 - 54 000
CENTRALE LYON - Post-doctoral researcher (12 months) Mathematics of energy-efficient AI: diffusion models on emerging hardware
CENTRALE LYON - Post-doctoral researcher (12 months) Mathematics of energy-efficient AI: diffusion models on emerging hardware

Centrale Lyon • Écully

Sur place
EUR 32 000 - 42 000
CENTRALE LYON - Post Doctoral Open and Flexible System-Level Evaluation Framework for Emerging AI Computing Architectures
CENTRALE LYON - Post Doctoral Open and Flexible System-Level Evaluation Framework for Emerging AI Computing Architectures

CENTRALE LYON • Écully

Hybride
EUR 38 000 - 54 000
HPC Performance Developer
HPC Performance Developer

Groupe EOLEN • Paris

Sur place
EUR 80 000 - 110 000
Access to Europe's largest supercomputers
Participation in international conferences
Cutting-edge R&D projects
Research Engineer Position - Java Developer - Corese Library
Research Engineer Position - Java Developer - Corese Library

Inria • France

Sur place
EUR 40 000 - 55 000
Partial reimbursement of public transport costs
7 weeks of annual leave
Possibility of teleworking
+2
Expert(e) en IA et calcul haute performance (H/F)
Expert(e) en IA et calcul haute performance (H/F)

CENTRE INFORMATIQUE NATIONAL DE L'ENSEIG • Montpellier

Sur place
EUR 23 000 - 33 000
Restauration
Ordinateur portable
Selon expérience ou grille fonctionn
CENTRALE LYON - Post Doctoral Open and Flexible System-Level Evaluation Framework for Emerging AI Computing Architectures
CENTRALE LYON - Post Doctoral Open and Flexible System-Level Evaluation Framework for Emerging AI Computing Architectures

ecolecentraledelyon • Écully

Hybride
EUR 40 000 - 56 000
ASIC Design Engineer - Hardware
ASIC Design Engineer - Hardware

Inria • Rennes

Sur place
EUR 27 369 - 32 783
Remboursement partiel des frais de交通
7 semaines de congés payés
Télétravail possible après 6 mois
+2
PhD Position F/M Frugal Distributed Training with Volatile Resources
PhD Position F/M Frugal Distributed Training with Volatile Resources

Inria • Valbonne

Sur place
EUR 23 357 - 27 978
Partial reimbursement of public transport costs
7 weeks of annual leave + 10 extra days off
Possibility of teleworking
+3