Cross-Tier and Distributed Caching and Data Management in Massively Distributed Systems

1000scholars

Rennes

In loco

EUR 42.000 - 62.000

Tempo pieno

9 giorni fa
Generatore di candidature

Una candidatura completa in un minuto — curriculum e lettera di presentazione personalizzati, pronti da inviare.

Supera i filtri ATS

Descrizione del lavoro

Inria Rennes, within the MAGELLAN team, seeks an engineer to study, implement and evaluate cross-tier caching and data management for hierarchical storage systems on the DXP project.

The role involves prototyping caching strategies, running experiments, reporting results, and collaborating with a PhD student and project partners including Amadeus and Inria teams COAST, CEDAR and MAGELLAN.

Competenze

  • Solid background in distributed systems and caching.
  • Experience building systems and tools.
  • Proficiency in Python and Java.
  • Knowledge of data management, storage and caching systems is advantageous.
  • Strong collaborative and networking skills.
  • Excellent written and oral English communication.

Mansioni

  • Study novel cross-tier and distributed caching strategies.
  • Prototype key caching strategies and data management techniques.
  • Run experiments and evaluate results.
  • Reporting, disseminating and presenting results.
  • Participate in project meetings with partners.

Conoscenze

distributed systems
building systems
Python
Java
data management
caching systems
collaboration
English communication

Strumenti

Python
Java

Descrizione del lavoro

Context

Financial and working environment.

This engineer position will be in the context of IPCEI-CIS (Important Project of Common European Interest – Next Generation Cloud Infrastructure and Services) DXP (Data Exchange Platform) project involving Amadeus and three Inria research teams (COAST, CEDAR and MAGELLAN). This project aims to design and develop an open-source management solution for a federated and distributed data exchange platform (DXP), operating in an open, scalable, and massively distributed environment (cloud-edge continuum). The position will be recruited and hosted at the Inria Center at Rennes University; and the work will be carried out within the MAGELLAN team in collaboration with other partners.

The position is for one year, with the possibility of an extension to 24 months.

Assignment

Context:

The ever-growing number of services and Internet of Things (IoT) devices has resulted in data being distributed across different locations (regions and countries) and different storage tiers. Additionally, data exhibits different usage patterns, including cold data (written once and never read), stream data (produced once and consumed by many), and hot data (written once and consumed by many). Furthermore, these data types have different performance and dependability requirements (e.g., low latency for data streams).

To ensure the reliability and improve the performance of data-intensive applications, data are either replicated or erasure-coded and distributed across different storage tiers, while frequently accessed data are stored on high-speed devices close to end users (i.e., cached). While much work has investigated data caching, data placement strategies (i.e., deciding what to cache), data movement, cache partitioning, cache eviction [1–8], and cost-efficient data redundancy techniques in caching systems [9], few efforts have focused holistic caching and data management when caches are distributed across heterogeneous platforms (from Edge to Cloud), utilize storage devices with varying performance and cost characteristics, and simultaneously serve diverse workloads, including traditional data services, serverless workflows, and data streaming.

The goal of this engineer position is to study, implement, and evaluate novel cross-tier and distributed caching strategies, alongside supporting data management techniques, for hierarchical multi-tier storage systems. The engineer will work closely with a PhD student on this topic.

References

[1] Asit Dan and Don Towsley. 1990. An Approximate Analysis of the LRU and FIFO Buffer Replacement Schemes. SIGMETRICS Perform. Eval. Rev. 18, 1 (apr 1990), 143–152. https://doi.org/10.1145/98460.98525

[2] Marek Chrobak and John Noga. 1999. LRU is better than FIFO. Algorithmica 23 (02 1999), 180–185. https://doi.org/10.1007/PL00009255

[3] Blankstein, Aaron, Siddhartha Sen, and Michael J. Freedman. “Hyperbolic caching: Flexible caching for web applications.” 2017 USENIX Annual Technical Conference (USENIX ATC 17). 2017.

[4] Cristian Ungureanu, Biplob Debnath, Stephen Rago, and Akshat Aranya. 2013. TBF: A memory-efficient replacement policy for flash- based caches. In 2013 IEEE 29th International Conference on Data Engineering (ICDE). 1117–1128. https://doi.org/10.1109/ICDE.2013.6544902

[5] Orcun Yildiz, Amelie Chi Zhou, Shadi Ibrahim. 2018. Improving the Effectiveness of Burst Buffers for Big Data Processing in HPC Systems with Eley. Future Generation Computer Systems, Volume 86, 2018, Pages 308-318, ISSN 0167-739X, https://doi.org/10.1016/j.future.2018.03.029

[6] G. Aupy, O. Beaumont and L. Eyraud-Dubois, "Sizing and Partitioning Strategies for Burst-Buffers to Reduce IO Contention," 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS), Rio de Janeiro, Brazil, 2019,

[7] ZHANG, Yazhuo, YANG, Juncheng, YUE, Yao, et al. {SIEVE} is simpler than {LRU}: an efficient {Turn-Key} eviction algorithm for web caches. In : 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). 2024. p. 1229-1246.

[8] Juncheng Yang, Ziming Mao, Yao Yue, and K. V. Rashmi. GL-Cache: Group-level learning for efficient and high-performance caching. FAST’23, pages 115–134, 2023.

[9] RASHMI, K. V., CHOWDHURY, Mosharaf, KOSAIAN, Jack, et al. {EC-Cache}:{Load-Balanced},{Low-Latency} cluster caching with online erasure coding. In : 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). 2016. p. 401-417.

Main activities
  • Study novel cross-tier and distributed caching strategies, alongside supporting data management techniques
  • Prototype key caching strategies and data management techniques
  • Run experiments and Evaluation of results
  • Reporting, disseminating and presenting results.
  • Participate in project meetings and discussions with other partners.
Skills
  • A solid background in the area of distributed systems
  • Experience with building systems and tools
  • Software development skills: Python and Java
  • Working experience in the areas of data management, storage and caching systems are advantageous
  • Good collaborative and networking skills
  • Excellent written and oral communication in English
Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Cross-Tier and Distributed Caching and Data Management in Massively Distributed Systems
Cross-Tier and Distributed Caching and Data Management in Massively Distributed Systems

Inria • Rennes

In loco
EUR 26.000 - 35.000
Public transport reimbursement
Annual leave and RTT days
Teleworking after 6 months
+3
Post-Doctoral Research Visit F/M Data management and job scheduling for Geo-distributed Workflows
Post-Doctoral Research Visit F/M Data management and job scheduling for Geo-distributed Workflows

1000scholars • Rennes

Ibrido
EUR 38.000 - 46.000
Edge-Cloud Data Management & Caching Engineer
Edge-Cloud Data Management & Caching Engineer

1000scholars • Rennes

In loco
EUR 42.000 - 62.000
PhD Position F/M Computational Storage and Its Application to Scientific Computing
PhD Position F/M Computational Storage and Its Application to Scientific Computing

1000scholars • Rennes

In loco
EUR 22.000 - 32.000
Coaching and mentoring
PhD Position F/M Computational Storage and Its Application to Scientific Computing
PhD Position F/M Computational Storage and Its Application to Scientific Computing

Inria • Rennes

In loco
EUR 17.000 - 21.000
Public transport refund
Leave: 7 weeks + RTT days
Teleworking after 6 months
+2
Cross-Tier Distributed Caching Engineer (Remote after 6mo)
Cross-Tier Distributed Caching Engineer (Remote after 6mo)

Inria • Rennes

In loco
EUR 26.000 - 35.000
Public transport reimbursement
Annual leave and RTT days
Teleworking after 6 months
+3
Making Caches Predictable: Bounded Lifetime Monitoring and Management for Real-Time Systems
Making Caches Predictable: Bounded Lifetime Monitoring and Management for Real-Time Systems

HiPEAC • Rennes

In loco
EUR 28.000 - 36.000
Engineer F/H - Adaptation of scientific codes to a task-based runtime system for exploiting high performance calibration and imaging pipelines in interferometric astronomy
Engineer F/H - Adaptation of scientific codes to a task-based runtime system for exploiting high performance calibration and imaging pipelines in interferometric astronomy

1000scholars • Talence

In loco
EUR 45.000 - 65.000
Engineer F/M - Development and optimization of a co-designed runtime system for running high performance machine learning chains
Engineer F/M - Development and optimization of a co-designed runtime system for running high performance machine learning chains

HiPEAC • Bordeaux

In loco
EUR 45.000 - 70.000
PhD Position F/M Predictable cache memories for Real-Time Embedded Multicore Systems
PhD Position F/M Predictable cache memories for Real-Time Embedded Multicore Systems

Inria, the French national research institute for the digital sciences • France

Ibrido
EUR 25.000 - 36.000
Transport partial
7 weeks annual leave
Teleworking after 6 months
+2