官术网_书友最值得收藏!

Introduction to Apache Spark

Apache Spark is an open source framework for processing large datasets stored in heterogeneous data stores in an efficient and fast way. Sophisticated analytical algorithms can be easily executed on these large datasets. Spark can execute a distributed program 100 times faster than MapReduce. As Spark is one of the fast-growing projects in the open source community, it provides a large number of libraries to its users.

We shall cover the following topics in this chapter:

  • A brief introduction to Spark
  • Spark architecture and the different languages that can be used for coding Spark applications
  • Spark components and how these components can be used together to solve a variety of use cases
  • A comparison between Spark and Hadoop
主站蜘蛛池模板: 淮北市| 墨脱县| 永春县| 腾冲县| 土默特左旗| 林周县| 汝州市| 龙南县| 宁海县| 都安| 张掖市| 乐至县| 汕尾市| 苏尼特左旗| 定远县| 商南县| 新化县| 饶河县| 佛坪县| 淳化县| 连江县| 左贡县| 介休市| 永春县| 滨州市| 千阳县| 宁德市| 宁津县| 绵竹市| 长武县| 隆德县| 成安县| 金湖县| 简阳市| 司法| 濉溪县| 同江市| 五常市| 晋江市| 阜新市| 阿拉善盟|