捕鱼小游戏可以玩的

書名： Natural Language Processing Fundamentals
作者名： Sohom Ghosh Dwight Gunning
本章字數： 138字
更新時間： 2021-06-11 13:42:31

Summary

In this chapter, you have learned about various types of data and ways to deal with unstructured text data. Text data is usually untidy and needs to be cleaned and pre-processed. Pre-processing steps mainly consist of tokenization, stemming, lemmatization, and stop-word removal. After pre-processing, features are extracted from texts using various methods, such as BoW and TF-IDF. This step converts unstructured text data into structured numeric data. New features are created from existing features using a technique called feature engineering. In the last part of the chapter, we explored various ways of visualizing text data, such as word clouds.

In the next chapter, you will learn how to develop machine learning models to classify texts using the features you have learned to extract in this chapter. Moreover, different sampling techniques and model evaluation parameters will be introduced.

官术网_书友最值得收藏!

Natural Language Processing Fundamentals

Summary