官术网_书友最值得收藏!

  • Big Data Analytics
  • Venkat Ankam
  • 430字
  • 2021-08-20 10:32:20

Chapter 1. Big Data Analytics at a 10,000-Foot View

The goal of this book is to familiarize you with tools and techniques using Apache Spark, with a focus on Hadoop deployments and tools used on the Hadoop platform. Most production implementations of Spark use Hadoop clusters and users are experiencing many integration challenges with a wide variety of tools used with Spark and Hadoop. This book will address the integration challenges faced with Hadoop Distributed File System (HDFS) and Yet Another Resource Negotiator (YARN) and explain the various tools used with Spark and Hadoop. This will also discuss all the Spark components—Spark Core, Spark SQL, DataFrames, Datasets, Spark Streaming, Structured Streaming, MLlib, GraphX, and SparkR and integration with analytics components such as Jupyter, Zeppelin, Hive, HBase, and dataflow tools such as NiFi. A real-time example of a recommendation system using MLlib will help us understand data science techniques.

In this chapter, we will approach Big Data analytics from a broad perspective and try to understand what tools and techniques are used on the Apache Hadoop and Apache Spark platforms.

Big Data analytics is the process of analyzing Big Data to provide past, current, and future statistics and useful insights that can be used to make better business decisions.

Big Data analytics is broadly classified into two major categories, data analytics and data science, which are interconnected disciplines. This chapter will explain the differences between data analytics and data science. Current industry definitions for data analytics and data science vary according to their use cases, but let's try to understand what they accomplish.

Data analytics focuses on the collection and interpretation of data, typically with a focus on past and present statistics. Data science, on the other hand, focuses on the future by performing explorative analytics to provide recommendations based on models identified by past and present data.

Figure 1.1 explains the difference between data analytics and data science with respect to time and value achieved. It also shows typical questions asked and tools and techniques used. Data analytics has mainly two types of analytics, descriptive analytics and diagnostic analytics. Data science has two types of analytics, predictive analytics and prescriptive analytics. The following diagram explains data science and data analytics:

Figure 1.1: Data analytics versus data science

The following table explains the differences with respect to processes, tools, techniques, skill sets, and outputs:

This chapter will cover the following topics:

  • Big Data analytics and the role of Hadoop and Spark
  • Big Data science and the role of Hadoop and Spark
  • Tools and techniques
  • Real-life use cases
主站蜘蛛池模板: 同心县| 张北县| 鄂托克旗| 井冈山市| 麦盖提县| 香格里拉县| 克什克腾旗| 巩义市| 公安县| 伊金霍洛旗| 商都县| 马公市| 襄城县| 辽宁省| 康定县| 连城县| 林周县| 乐亭县| 古浪县| 乃东县| 乌海市| 天镇县| 遂宁市| 彭阳县| 乌兰县| 黎平县| 石城县| 开江县| 滨州市| 峨边| 金沙县| 东明县| 东城区| 鸡西市| 香港| 图木舒克市| 崇礼县| 乐安县| 武平县| 白朗县| 信阳市|