Preprint
Machine Learning

A survey on data storage and placement methodologies for Cloud-Big Data ecosystem

Somnath Mazumdar(Simula Research Laboratory), Daniel Seybold(Universität Ulm), Kyriakos Kritikos(FORTH Institute of Computer Science), Yiannis Verginadis(Institute of Communication and Computer Systems)
February 11, 2019Journal Of Big Data123 citations

123

Citations

1

Influential Citations

Journal Of Big Data

Venue

2019

Year

Abstract

Currently, the data to be explored and exploited by computing systems increases at an exponential rate. The massive amount of data or so-called “Big Data” put pressure on existing technologies for providing scalable, fast and efficient support. Recent applications and the current user support from multi-domain computing, assisted in migrating from data-centric to knowledge-centric computing. However, it remains a challenge to optimally store and place or migrate such huge data sets across data centers (DCs). In particular, due to the frequent change of application and DC behaviour (i.e., resources or latencies), data access or usage patterns need to be analyzed as well. Primarily, the main objective is to find a better data storage location that improves the overall data placement cost as well as the application performance (such as throughput). In this survey paper, we are providing a state of the art overview of Cloud-centric Big Data placement together with the data storage methodologies. It is an attempt to highlight the actual correlation between these two in terms of better supporting Big Data management. Our focus is on management aspects which are seen under the prism of non-functional properties. In the end, the readers can appreciate the deep analysis of respective technologies related to the management of Big Data and be guided towards their selection in the context of satisfying their non-functional application requirements. Furthermore, challenges are supplied highlighting the current gaps in Big Data management marking down the way it needs to evolve in the near future.

Analysis

Why This Paper Matters

This survey addresses a critical bottleneck in the Cloud-Big Data ecosystem: the optimal storage and placement of exponentially growing datasets across data centers. As applications shift from data-centric to knowledge-centric computing, the need for scalable, fast, and efficient data management becomes paramount. The paper uniquely focuses on non-functional properties such as cost and throughput, which are often overlooked in favor of functional requirements. By providing a structured overview of existing methodologies, it helps practitioners navigate the complex landscape of cloud-based Big Data storage and placement, directly impacting application performance and operational costs.

Technical Contributions

The paper's main contributions are:

  • Comprehensive survey: It systematically reviews cloud-centric Big Data placement and storage methodologies, covering a wide range of techniques.
  • Correlation analysis: It explicitly highlights the interplay between data storage and placement strategies for improved Big Data management.
  • Non-functional property focus: It evaluates technologies through the lens of non-functional properties (e.g., latency, throughput, cost), which are crucial for real-world deployments.
  • Guidance for selection: It provides readers with criteria to choose appropriate technologies based on their non-functional application requirements.
  • Challenge identification: It outlines current gaps and future evolution directions in Big Data management, such as handling dynamic application and data center behavior.

Results

As a survey, the paper does not present new experimental results or quantitative metrics. Instead, it synthesizes findings from 123 cited works to offer a qualitative analysis of existing technologies. The key outcome is a structured understanding of how different storage and placement methods affect non-functional properties, enabling informed decision-making. The paper also identifies open challenges, such as the need for adaptive strategies that respond to frequent changes in application and data center behavior.

Significance

This survey serves as a valuable reference for AI practitioners and cloud architects dealing with large-scale data. By focusing on non-functional properties, it bridges the gap between theoretical data management techniques and practical deployment considerations. The identified challenges point to future research directions, such as developing more adaptive and cost-aware placement algorithms. For the broader AI field, efficient data storage and placement directly impact the performance of data-intensive machine learning pipelines, making this work relevant to optimizing end-to-end AI systems in cloud environments.