官术网_书友最值得收藏!

Machine learning algorithms

In this section, we review algorithms that are needed for machine learning, and introduce machine learning libraries including Spark's MLlib and IBM's SystemML, then we discuss their integration with Apache Spark.

After reading this section, readers will become familiar with various machine learning libraries including Spark's MLlib, and know how to make them ready for machine learning.

To complete a Machine Learning project, data scientists often employ some classification or regression algorithms to develop and evaluate predictive models, which are readily available in some Machine Learning tools like R or MatLab. To complete a machine learning project, besides data sets and computing platforms, these machine learning libraries, as collections of machine learning algorithms, are necessary.

For example, the strength and depth of the popular R mainly comes from the various algorithms that are readily provided for the use of Machine Learning professionals. The total number of R packages is over 1000. Data scientists do not need all of them, but do need some packages to:

  • Load data, with packages like RODBC or RMySQL
  • Manipulate data, with packages like stringr or lubridate
  • Visualize data, with packages like ggplot2 or leaflet
  • Model data, with packages like Random Forest or survival
  • Report results, with packages like shiny or markdown

According to a recent ComputerWorld survey, the most downloaded R packages are:

主站蜘蛛池模板: 焦作市| 日喀则市| 县级市| 海口市| 新营市| 昔阳县| 兰坪| 诸暨市| 西吉县| 武川县| 凤台县| 招远市| 呈贡县| 保康县| 宁乡县| 筠连县| 榆林市| 黄陵县| 兴隆县| 兴隆县| 马尔康县| 临夏市| 岑巩县| 苍溪县| 淄博市| 历史| 德昌县| 杭锦后旗| 嘉鱼县| 西丰县| 中超| 八宿县| 桐城市| 银川市| 荃湾区| 洞口县| 曲阜市| 江油市| 武定县| 杭锦旗| 全椒县|