An application made for this job — a tailored resume and cover letter that speak straight to the posting.
KerData at Inria Rennes invites applications for a PhD thesis in computational storage for scientific computing and AI. The project emphasises ontology development, empirical evaluation, and scalable modeling across distributed platforms, with strong collaboration with international partners.
You will join a small, international team and work within top-tier HPC research environments, benefiting from personalized coaching and access to world-class hardware and software stacks for rigorous
Supervisory Team:
The thesis will be hosted by the KerData team at the Inria research center of Rennes. Rennes is the capital city of Britanny, in the western part of France. It is easy to reach thanks to the high-speed train line to Paris. Rennes is a dynamic, lively city and a major center for higher education and research: 25% of its population are students.
This thesis will include collaborations with international partners from Germany, thus research visits to and from the collaborator's teams are expected.
KerData is a human-sized team currently comprising 5 permanent researchers, 2 engineers and 6 PhD students. You will work in a caring environment, offering a good work-life balance.
KerData is leading multiple projects in top-level national and international collaborative environments such as within the Joint-Laboratory on Extreme-Scale Computing: https://jlesc.github.io. Our team has active collaboration with high-profile academic institutions all around the world (including the USA, Spain, Germany or Japan) and with industry.
Our team strongly favors experimental research, validated by implementation and experimentation of software prototypes with real-world applications on real-world platforms including some of the most powerful supercomputers worldwide.
The KerData team is committed to personalized advising and coaching, to help PhD candidates train and grow in all directions that are critical in the process of becoming successful researchers.
Check our website for more about the KerData team here: https://team.inria.fr/kerdata/
Computational storage [15, 16] is a promising technology to improve the efficiency and performance of various workloads, for example in scientific computing, machine learning (ML), and artificial intelligence (AI). As the volume of generated data continues to grow exponentially [5,9] conventional compute and storage architectures are increasingly constrained by large data movements between storage/memory and compute resources. Computational storage can eliminate many of these data movements by co-locating compute capabilities along with storage locations allowing to offload the typically much smaller (sub)programs [1, 12, 14]. While the concept is well established in literature, computational storage devices are not widely commercially available or deployed in data centers today. A key challenge is that computational storage encourages domain-specialization for highest efficiency while economic factors encourage commoditization of products catering to broad markets. A second challenge is that both for legacy applications as well as emerging applications such as ML/AI it remains an open research question how to program and orchestrate across distributed platforms with computational storage capabilities.
Recent advancements in programming models and software portability on the one hand, and reconfigurable hardware and domain-specific hardware design on the other [5, 7, 11], suggest that a modular approach that identifies common building blocks across domain boundaries might hold the key to both aforementioned challenges. Computational storage research exists on accelerating specific workloads or applications [5, 11, 14] as well as on emulating computational storage devices but a systematic study focusing on scientific computing workloads and modelling of suitable architectures and data distribution strategies is missing.
This project aims to advance the research on computational storage for scientific computing and artificial intelligence applications. It will investigate mechanisms to formalize, capture, model and evaluate computational storage in distributed environments. The project is structured into three primary objectives:
To explore how computational storage can aid workloads in scientific computing and artificial intelligence, we will build upon previous work and active research of the members of the supervisory team in Germany and France.
For objective A, the research methodology centers on analysing real-world scientific computing use cases in close exchange with domain scientists to identify computational storage opportunities. This work will establish ontologies and taxonomies for distributed computational storage systems from multiple angles across multiple domains. This work will be complementary to collaborations with the German Climate Computing Center (DKRZ) and the Parallel Computing and I/O group at Otto von Guericke University Magdeburg. Several surveys have taken snapshots of the state of the art of computational storage [12, 14] and its precursors or related concepts (e.g., active storage, near-data processing, processing in memory). The outcome of this aim will be a state of the art survey focusing on the applicability for scientific computing in the HPC, to cloud and edge computing continuum.
Objective B is to establish suitable testbeds and modeling environments to study computational storage at scale. The work will build upon existing research of the KerData team to leverage system simulation to model large scale distributed systems [13] in addition to empirical platforms leveraging hardware emulation [7, 10], as well as realistic software stacks and physical hardware within Grid5000, Slices-FR or Chameleon Cloud and domain-specific experimental platforms together with collaborators from the different domain sciences. The targeted outcome of Objective B are proof-of-concept environments to run real-world computational workflows leveraging real operating system, middleware and device APIs.
Objective C is to develop a methodology to faithfully model performance extrapolation to large-scale deployment scenarios will evaluate and validate the impact achievable through common computational storage building blocks. The methodology aims to study, for example, task and data placement strategies based on the needs of real-world computational science use cases but extrapolated to system scales not deployed in state of the art data centers today. The outcome will be a methodological framework to support middleware and domain-specific computational storage device development.
Required Competence:
Appreciated: