官术网_书友最值得收藏!

Summary

In this chapter, we introduced the concept of supervised machine learning, along with a number of use cases, including the automation of manual tasks such as identifying hairstyles from the 1960s and 1980s. In this introduction, we encountered the concept of labeled datasets and the process of mapping one information set (the input data or features) to the corresponding labels.

We took a practical approach to the process of loading and cleaning data using Jupyter notebooks and the extremely powerful pandas library. Note that this chapter has only covered a small fraction of the functionality within pandas, and that an entire book could be dedicated to the library itself. It is recommended that you become familiar with reading the pandas documentation and continue to develop your pandas skills through practice.

The final section of this chapter covered a number of data quality issues that need to be considered to develop a high-performing supervised learning model, including missing data, class imbalance, and low sample sizes. We discussed a number of options for managing such issues and emphasized the importance of checking these mitigations against the performance of the model.

In the next chapter, we will extend upon the data cleaning process that we covered and will investigate the data exploration and visualization process. Data exploration is a critical aspect of any machine learning solution, as without a comprehensive knowledge of the dataset, it would be almost impossible to model the information provided.

主站蜘蛛池模板: 额尔古纳市| 焉耆| 金沙县| 霸州市| 石棉县| 伊金霍洛旗| 泽普县| 台州市| 阳原县| 牡丹江市| 抚顺市| 广水市| 中江县| 邯郸县| 古丈县| 黑水县| 刚察县| 乌恰县| 博罗县| 清涧县| 前郭尔| 德令哈市| 鄂托克前旗| 屏东县| 冕宁县| 永善县| 奉节县| 济源市| 苏尼特右旗| 兰坪| 四会市| 吉林市| 日土县| 东山县| 博客| 涪陵区| 江城| 太湖县| 宾川县| 七台河市| 吕梁市|