官术网_书友最值得收藏!

Summary

In this chapter, we covered how to set up Spark locally on our own computer as well as in the cloud as a cluster running on Amazon EC2. You learned how to run Spark on top of Amazon's Elastic Map Reduce (EMR). You also learned how to use Google Compute Engine's Spark Service to create a cluster and run a simple job. We discussed the basics of Spark's programming model and API using the interactive Scala console, and we wrote the same basic Spark program in Scala, Java, R, and Python. We also compared the performance metrics of Hadoop versus Spark for different machine learning algorithms as well as SORT benchmark tests.

In the next chapter, we will consider how to go about using Spark to create a machine learning system.

主站蜘蛛池模板: 梁平县| 桃园市| 星子县| 仙桃市| 奉化市| 东平县| 竹溪县| 上犹县| 太保市| 乡城县| 辉南县| 桃园县| 平塘县| 平湖市| 丹巴县| 云霄县| 柳江县| 东乡县| 满城县| 孟州市| 兴海县| 于都县| 那曲县| 桦甸市| 巴南区| 古浪县| 项城市| 永德县| 凉城县| 沂源县| 松潘县| 柘荣县| 赤水市| 临夏县| 江阴市| 牡丹江市| 湘乡市| 贵溪市| 河曲县| 克什克腾旗| 奉化市|