官术网_书友最值得收藏!

What this book covers

Chapter 1, The Need for Data Lake, helps you understand what Data Lake is, its architecture and key components, and the business contexts where Data Lake can be successfully deployed. You will also learn the limitations of the traditional data architectures and how Data Lake addresses some of these inadequacies and provides significant benefits.

Chapter 2, Data Intake, helps you understand the Intake Tier in detail where we will explore the process of obtaining huge volumes of data into Data Lake. You will learn the technology perspective of the various External Data Sources and Hadoop-based data transfer mechanisms to pull or push data into Data Lake.

Chapter 3, Data Integration, Quality, and Enrichment, explores the processes that are performed on vast quantities of data in the Management Tier. You will get a deeper understanding of the key technology aspects and components such as profiling, validation, integration, cleansing, standardization, and enrichment using Hadoop ecosystem components.

Chapter 4, Data Discovery and Consumption, helps you understand how data can be discovered, packaged, and provisioned, for it to be consumed by the downstream systems. You will learn the key technology aspects, architectural guidance and tools for data discovery, and data provisioning functionalities.

Chapter 5, Data Governance, explores the details, need, and utility of data governance in a Data Lake environment. You will learn how to deal with metadata management, lineage tracking, data lifecycle management to govern the usability, security, integrity, and availability of the data through the data governance processes applied on the data in Data Lake. This chapter also explores how the current Data Lake can evolve in a futuristic setting.

主站蜘蛛池模板: 固始县| 博湖县| 思南县| 建瓯市| 温州市| 股票| 封丘县| 手游| 三门峡市| 徐汇区| 甘孜县| 无极县| 旺苍县| 镇江市| 天柱县| 辽宁省| 呈贡县| 中阳县| 同心县| 和林格尔县| 都匀市| 夹江县| 峨眉山市| 乌恰县| 黄平县| 京山县| 边坝县| 辽宁省| 台南市| 两当县| 鄂尔多斯市| 介休市| 车致| 莲花县| 舒城县| 灵川县| 木兰县| 河间市| 塔城市| 香格里拉县| 田林县|