A Big Data Lake for Multilevel Streaming Analytics

Liu, Ruoran; Isah, Haruna; Zulkernine, Farhana

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:2009.12415 (cs)

[Submitted on 25 Sep 2020]

Title:A Big Data Lake for Multilevel Streaming Analytics

Authors:Ruoran Liu, Haruna Isah, Farhana Zulkernine

View PDF

Abstract:Large organizations are seeking to create new architectures and scalable platforms to effectively handle data management challenges due to the explosive nature of data rarely seen in the past. These data management challenges are largely posed by the availability of streaming data at high velocity from various sources in multiple formats. The changes in data paradigm have led to the emergence of new data analytics and management architecture. This paper focuses on storing high volume, velocity and variety data in the raw formats in a data storage architecture called a data lake. First, we present our study on the limitations of traditional data warehouses in handling recent changes in data paradigms. We discuss and compare different open source and commercial platforms that can be used to develop a data lake. We then describe our end-to-end data lake design and implementation approach using the Hadoop Distributed File System (HDFS) on the Hadoop Data Platform (HDP). Finally, we present a real-world data lake development use case for data stream ingestion, staging, and multilevel streaming analytics which combines structured and unstructured data. This study can serve as a guide for individuals or organizations planning to implement a data lake solution for their use cases.

Comments:	6 pages
Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Databases (cs.DB); Machine Learning (cs.LG)
Cite as:	arXiv:2009.12415 [cs.DC]
	(or arXiv:2009.12415v1 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.2009.12415

Submission history

From: Haruna Isah [view email]
[v1] Fri, 25 Sep 2020 19:57:21 UTC (477 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:A Big Data Lake for Multilevel Streaming Analytics

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:A Big Data Lake for Multilevel Streaming Analytics

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators