- Apache Spark Machine Learning Blueprints
- Alex Liu
- 224字
- 2021-07-16 10:39:50
Chapter 2. Data Preparation for Spark ML
Machine learning professionals and data scientists often spend 70% or 80% of their time preparing data for their machine learning projects. Data preparation can be very hard work, but it is necessary and extremely important as it affects everything to follow. Therefore, in this chapter, we will cover all the necessary data preparation parts for our machine learning, which often runs from data accessing, data cleaning, datasets joining, and then to feature development so as to get our datasets ready to develop ML models on Spark. Specifically, we will discuss the following six data preparation tasks mentioned before and then end our chapter with a discussion of repeatability and automation:
- Accessing and loading datasets
- Publicly available datasets for ML
- Loading datasets into Spark easily
- Exploring and visualizing data with Spark
- Data cleaning
- Dealing with missing cases and incompleteness
- Data cleaning on Spark
- Data cleaning made easy
- Identity matching
- Dealing with identity issues
- Data matching on Spark
- Data matching made better
- Data reorganizing
- Data reorganizing tasks
- Data reorganizing on Spark
- Data reorganizing made easy
- Joining data
- Spark SQL to join datasets
- Joining data with Spark SQL
- Joining data made easy
- Feature extraction
- Feature extraction challenges
- Feature extraction on Spark
- Feature extraction made easy
- Repeatability and automation
- Dataset preprocessing workflows
- Spark pipelines for preprocessing
- Dataset preprocessing automation
推薦閱讀
- Java編程全能詞典
- 面向STEM的mBlock智能機(jī)器人創(chuàng)新課程
- Dreamweaver CS3網(wǎng)頁設(shè)計(jì)50例
- Hands-On Cloud Solutions with Azure
- 條碼技術(shù)及應(yīng)用
- 計(jì)算機(jī)網(wǎng)絡(luò)應(yīng)用基礎(chǔ)
- JMAG電機(jī)電磁仿真分析與實(shí)例解析
- 網(wǎng)絡(luò)綜合布線技術(shù)
- STM32嵌入式微控制器快速上手
- JSF2和RichFaces4使用指南
- 可編程序控制器應(yīng)用實(shí)訓(xùn)(三菱機(jī)型)
- Visual FoxPro程序設(shè)計(jì)
- 從零開始學(xué)Java Web開發(fā)
- Machine Learning Algorithms(Second Edition)
- Redash v5 Quick Start Guide