PhD Position F/M Computational Storage and Its Application to Scientific Computing

Inria

Rennes

Sur place

EUR 17 000 - 21 000

Plein temps

Il y a 10 jours
Générateur de candidature

N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Avantages offerts par ce poste

Public transport refund
Leave: 7 weeks + RTT days
Teleworking after 6 months
Equipment provided
Social, cultural and sports events

Résumé du poste

Inria is seeking a PhD candidate to join the KerData team in Rennes. The project focuses on computational storage and its applications to scientific computing and AI, with collaborations across European partners.

You will pursue three objectives: survey ontologies, build empirical testbeds, and model performance for large-scale deployments. The position offers modern research facilities and close mentoring by senior researchers, with opportunities to publish in top venues and contribute to open

Qualifications

  • Excellent academic record in computer science.
  • Knowledge on distributed systems and data management.
  • Familiarity with containerized environments (Docker, Podman, Apptainer, Kubernetes).
  • Strong motivation to conduct high‑quality research and publish.
  • Excellent English communication skills, both written and oral.

Responsabilités

  • Carry out doctoral research on computational storage for scientific computing.
  • Develop ontologies, taxonomies and modeling methods for distributed storage systems.
  • Collaborate with international partners and within the KerData team.
  • Prototype software and run experiments on HPC platforms and datasets.

Connaissances

Academic excellence in CS
Distributed systems
Data management systems
Containerized environments
Research communication (English)
Team collaboration

Formation

PhD in Computer Science

Outils

Docker
Podman
Apptainer
Kubernetes

Description du poste

PhD Position F/M Computational Storage and Its Application to Scientific Computing

Fonction : Doctorant

The Inria center at the University of Rennes is one of eight Inria centers and has more than thirty research teams. The Inria center is a major and recognized player in the field of digital sciences. It is at the heart of a rich ecosystem of R&D and innovation, including highly innovative SMEs, large industrial groups, competitiveness clusters, research and higher education institutions, centers of excellence, and technological research institutes.

Supervisory Team:
  • Jakob Luettgau, PhD (Inria, France)
Location and Mobility

The thesis will be hosted by the KerData team at the Inria research center of Rennes. Rennes is the capital city of Britanny, in the western part of France. It is easy to reach thanks to the high-speed train line to Paris. Rennes is a dynamic, lively city and a major center for higher education and research: 25% of its population are students.

This thesis will include collaborations with international partners from Germany, thus research visits to and from the collaborator's teams are expected.

The KerData team in a nutshell for candidates
  • KerData is a human-sized team currently comprising 5 permanent researchers, 2 engineers and 6 PhD students. You will work in a caring environment, offering a good work-life balance.
  • KerData is leading multiple projects in top-level national and international collaborative environments such as within the Joint-Laboratory on Extreme-Scale Computing: https://jlesc.github.io. Our team has active collaboration with high-profile academic institutions all around the world (including the USA, Spain, Germany or Japan) and with industry.
  • Our team strongly favors experimental research, validated by implementation and experimentation of software prototypes with real-world applications on real-world platforms including some of the most powerful supercomputers worldwide.
  • The KerData team is committed to personalized advising and coaching, to help PhD candidates train and grow in all directions that are critical in the process of becoming successful researchers.

Computational storage [15, 16] is a promising technology to improve the efficiency and performance of various workloads, for example in scientific computing, machine learning (ML), and artificial intelligence (AI). As the volume of generated data continues to grow exponentially [5,9] conventional compute and storage architectures are increasingly constrained by large data movements between storage/memory and compute resources. Computational storage can eliminate many of these data movements by co-locating compute capabilities along with storage locations allowing to offload the typically much smaller (sub)programs [1, 12, 14]. While the concept is well established in literature, computational storage devices are not widely commercially available or deployed in data centers today. A key challenge is that computational storage encourages domain-specialization for highest efficiency while economic factors encourage commoditization of products catering to broad markets. A second challenge is that both for legacy applications as well as emerging applications such as ML/AI it remains an open research question how to program and orchestrate across distributed platforms with computational storage capabilities.

Recent advancements in programming models and software portability on the one hand, and reconfigurable hardware and domain-specific hardware design on the other [5, 7, 11], suggest that a modular approach that identifies common building blocks across domain boundaries might hold the key to both aforementioned challenges. Computational storage research exists on accelerating specific workloads or applications [5, 11, 14] as well as on emulating computational storage devices but a systematic study focusing on scientific computing workloads and modelling of suitable architectures and data distribution strategies is missing.

Research Objectives

This project aims to advance the research on computational storage for scientific computing and artificial intelligence applications. It will investigate mechanisms to formalize, capture, model and evaluate computational storage in distributed environments. The project is structured into three primary objectives:

  • Objective A: realize a survey and define ontologies and taxonomies for distributed computational storage systems from multiple angles across multiple domains: the application perspective (programming paradigm, data flows), the middleware perspective (workflow orchestration, resource management and data placement), and the system perspective (e.g., hardware, storage, compute)
  • Objective B: establish the technical foundations to empirically assess hardware performance and decompose existing workflows and task characteristics on the one hand, as well as implement suitable computational storage testbeds (e.g. simulated [2,3,4], emulated [7, 10], and on realistic software stacks and physical hardware [6, 8, 9, 11])
  • Objective C: develop a methodology to faithfully model performance extrapolations to large-scale deployment scenarios as a means to identify and validate common computational storage building blocks
Envisioned Approach

To explore how computational storage can aid workloads in scientific computing and artificial intelligence, we will build upon previous work and active research of the members of the supervisory team in Germany and France.

For objective A, the research methodology centers on analysing real-world scientific computing use cases in close exchange with domain scientists to identify computational storage opportunities. This work will establish ontologies and taxonomies for distributed computational storage systems from multiple angles across multiple domains. This work will be complementary to collaborations with the German Climate Computing Center (DKRZ) and the Parallel Computing and I/O group at Otto von Guericke University Magdeburg. Several surveys have taken snapshots of the state of the art of computational storage [12, 14] and its precursors or related concepts (e.g., active storage, near-data processing, processing in memory). The outcome of this aim will be a state of the art survey focusing on the applicability for scientific computing in the HPC, to cloud and edge computing continuum.

Objective B is to establish suitable testbeds and modeling environments to study computational storage at scale. The work will build upon existing research of the KerData team to leverage system simulation to model large scale distributed systems [13] in addition to empirical platforms leveraging hardware emulation [7, 10], as well as realistic software stacks and physical hardware within Grid5000, Slices-FR or Chameleon Cloud and domain-specific experimental platforms together with collaborators from the different domain sciences. The targeted outcome of Objective B are proof-of-concept environments to run real-world computational workflows leveraging real operating system, middleware and device APIs.

Objective C is to develop a methodology to faithfully model performance extrapolation to large-scale deployment scenarios will evaluate and validate the impact achievable through common computational storage building blocks. The methodology aims to study, for example, task and data placement strategies based on the needs of real-world computational science use cases but extrapolated to system scales not deployed in state of the art data centers today. The outcome will be a methodological framework to support middleware and domain-specific computational storage device development.

References

[1] Antonio Barbalace and Jaeyoung Do. 2021. Computational Storage: Where Are We Today? (2021). Retrieved from https://www.research.ed.ac.uk/en/publications/computational-storage-where-are-we-today

[2] H. Casanova. 2001. Simgrid: a toolkit for the simulation of application scheduling. In Proceedings First IEEE/ACM International Symposium on Cluster Computing and the Grid, May 2001. 430–437. https://doi.org/10.1109/CCGRID.2001.923223

[3] Henri Casanova, Rafael Ferreira da Silva, Ryan Tanaka, Suraj Pandey, Gautam Jethwani, William Koch, Spencer Albrecht, James Oeth, and Frédéric Suter. 2020. Developing accurate and scalable simulators of production workflow management systems with WRENCH. Future Generation Computer Systems 112, (November 2020), 162–175. https://doi.org/10.1016/j.future.2020.05.030

[4] Henri Casanova, Arnaud Giersch, Arnaud Legrand, Martin Quinson, and Frédéric Suter. 2014. Versatile, Scalable, and Accurate Simulation of Distributed Applications and Platforms. Journal of Parallel and Distributed Computing 74, 10 (June 2014), 2899. https://doi.org/10.1016/j.jpdc.2014.06.008

[5] Mehdi Hassanpour, Marc Riera, and Antonio González. 2021. A Survey of Near-Data Processing Architectures for Neural Networks. https://doi.org/10.48550/arXiv.2112.12630

[6] Niclas Hedam, Morten Tychsen Clausen, Philippe Bonnet, Sangjin Lee, and Ken Friis Larsen. 2023. Delilah: eBPF-offload on Computational Storage. In Proceedings of the 19th International Workshop on Data Management on New Hardware (DaMoN ’23), June 18, 2023. Association for Computing Machinery, New York, NY, USA, 70–76. https://doi.org/10.1145/3592980.3595319

[7] Sang-Hoon Kim, Jaehoon Shim, Euidong Lee, Seongyeop Jeong, Ilkueon Kang, and Jin-Soo Kim. 2023. NVMeVirt: A Versatile Software-defined Virtual NVMe Device. 2023. 379–394. Retrieved September 22, 2025 from https://www.usenix.org/conference/fast23/presentation/kim-sang-hoon

[8] Jaewook Kwak, Sangjin Lee, Kibin Park, Jinwoo Jeong, and Yong Ho Song. 2020. Cosmos+ OpenSSD: Rapid Prototype for Flash Storage Systems. ACM Trans. Storage 16, 3 (August 2020), 1–35. https://doi.org/10.1145/3385073

[9] Joo Hwan Lee, Hui Zhang, Veronica Lagrange, Praveen Krishnamoorthy, Xiaodong Zhao, and Yang Seok Ki. 2020. SmartSSD: FPGA Accelerated Near-Storage Data Analytics on SSD. IEEE Computer Architecture Letters 19, 2 (July 2020), 110–113. https://doi.org/10.1109/LCA.2020.3009347

[10] Huaicheng Li, Mingzhe Hao, Michael Hao Tong, Swaminatahan Sundararaman, and Haryadi S Gunawi. The CASE of FEMU: Cheap, Accurate, Scalable and Extensible Flash Emulator.

[11] Leibo Liu, Jianfeng Zhu, Zhaoshi Li, Yanan Lu, Yangdong Deng, Jie Han, Shouyi Yin, and Shaojun Wei. 2019. A Survey of Coarse-Grained Reconfigurable Architecture and Design: Taxonomy, Challenges, and Applications. ACM Comput. Surv. 52, 6 (October 2019), 118:1-118:39. https://doi.org/10.1145/3357375

[12] Corne Lukken and Animesh Trivedi. 2021. Past, Present and Future of Computational Storage: A Survey. https://doi.org/10.48550/arXiv.2112.09691

[13] Julien Monniot, François Tessier, Henri Casanova, and Gabriel Antoniu. 2024. Simulation of Large-Scale HPC Storage Systems: Challenges and Methodologies. December 18, 2024. 1. Retrieved January 4, 2025 from https://inria.hal.science/hal-04784808

[14] Siqi Zhang, Na Yi, and Yi Ma. 2024. A Survey of Computation Offloading with Task Types. https://doi.org/10.48550/arXiv.2401.01017

[15] 2025. NVM Express Computational Programs Command Set Specification, Revision 1.1.

[16] 2026. SNIA Computational Storage Architecture and Programming Model 1.2. Retrieved June 30, 2026 from https://www.snia.org/sites/default/files/technical-work/computational/release/SNIA-Computational-Storage-Architecture-and-Programming-Model-1.2.pdf

Required Competence:

  • An excellent academic record in computer science courses
  • Knowledge on distributed systems and data management systems
  • Familiarity with containerized environments (e.g., Docker, Podman, Apptainer, Kubernetes)
  • Ability and motivation to conduct high-quality research, including publishing the results in relevant venues
  • Very good communication skills in oral and written English
  • Open-mindedness, strong integration skills and team spirit

Appreciated:

  • Knowledge on machine learning and data analysis methods
  • Professional experience in the areas of HPC and Big Data management
Avantages
  • Partial reimbursement of public transport costs
  • Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
  • Possibility of teleworking (after 6 months of employment) and flexible organization of working hours
  • Professional equipment available (videoconferencing, loan of computer equipment, etc.)
  • Social, cultural and sports events and activities
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Cross-Tier and Distributed Caching and Data Management in Massively Distributed Systems
Cross-Tier and Distributed Caching and Data Management in Massively Distributed Systems

Inria • Rennes

Sur place
EUR 26 000 - 35 000
Public transport reimbursement
Annual leave and RTT days
Teleworking after 6 months
+3
PhD Position F/M Frugal Distributed Training with Volatile Resources
PhD Position F/M Frugal Distributed Training with Volatile Resources

Inria • Valbonne

Sur place
Partial reimbursement of public transport costs
7 weeks of annual leave + 10 extra days off
Possibility of teleworking
+3
PhD Position F/M Pretrained models of multimodal neuroimaging for predicting individual cognition
PhD Position F/M Pretrained models of multimodal neuroimaging for predicting individual cognition

Inria • Palaiseau

Hybride
EUR 18 000 - 30 000
Remboursement des frais de transport
7 semaines de congés + RTT
Télétravail possible et organisation d
+2
Kokkos supporting for complex data discretization and unstructured meshes
Kokkos supporting for complex data discretization and unstructured meshes

Inria • Palaiseau

Sur place
EUR 40 000 - 65 000
Partial reimbursement of public transport costs
7 weeks of annual leave
Possibility of teleworking after 6 months
+2
PhD in Computational Storage for Scientific Computing | Flexible Hours
PhD in Computational Storage for Scientific Computing | Flexible Hours

Inria • Rennes

Sur place
EUR 17 000 - 21 000
Public transport refund
Leave: 7 weeks + RTT days
Teleworking after 6 months
+2
HPC Performance Developer
HPC Performance Developer

Groupe EOLEN • Paris

Sur place
EUR 80 000 - 110 000
Access to Europe's largest supercomputers
Participation in international conferences
Cutting-edge R&D projects
Doctorant F/H Étude des plateformes dédiées à l'expérimentation des systèmes distribués: usages et exploitation
Doctorant F/H Étude des plateformes dédiées à l'expérimentation des systèmes distribués: usages et exploitation

Inria • Grenoble

Sur place
EUR 21 000 - 27 000
Restauration subventionnée
Transports publics remboursés
Congés généreux (7 semaines + RTT)
+3
PhD Position F/M Trustable Machine Learning : Analysis and Verification of Soft Automata
PhD Position F/M Trustable Machine Learning : Analysis and Verification of Soft Automata

Inria • Rennes

Sur place
EUR 23 000 - 28 000
Remboursement partiel des transports
Congés annuels + RTT
Télétravail possible après 6 mois
+1
Post-doctoral researcher in mechanistic interpretability and diversity of representations
Post-doctoral researcher in mechanistic interpretability and diversity of representations

IMT Atlantique • France

Sur place
EUR 36 000 - 45 000
Partial working from home
Public transport paid for
Doctorant Systèmes d'Exploitation pour Base Cerf-Valeur Rapide H/F
Doctorant Systèmes d'Exploitation pour Base Cerf-Valeur Rapide H/F

INRIA • Villé

Hybride
EUR 23 000 - 28 000
Restauration subventionnée
Transports publics remboursés
Congés: 7 semaines + RTT
+2