官术网_书友最值得收藏!

Semantics and topic modeling

Gensim is famous for its powerful semantic and topic modeling algorithms. Topic modeling is a typical text mining task of discovering the hidden semantic structures in a document. Semantic structure in plain English is the distribution of word occurrences. It is obviously an unsupervised learning task. What we need to do is to feed in plain text and let the model figure out the abstract "topics". We will study topic modeling in detail in Chapter 3, Mining the 20 Newsgroups Dataset with Clustering and Topic Modeling Algorithms.

In addition to robust semantic modeling methods, gensim also provides the following functionalities:

  • Word embedding: Also known as word vectorization, this is an innovative way to represent words while preserving words' co-occurrence features. We will study word embedding in detail in Chapter 10, Machine Learning Best Practices.
  • Similarity querying: This functionality retrieves objects that are similar to the given query object. It's a feature built on top of word embedding.
  • Distributed computingThis functionality makes it possible to efficiently learn from millions of documents.

Last but not least, as mentioned in the first chapter, scikit-learn is the main package we use throughout this entire book. Luckily, it provides all text processing features we need, such as tokenization, besides comprehensive machine learning functionalities. Plus, it comes with a built-in loader for the 20 newsgroups dataset.

Now that the tools are available and properly installed, what about the data?

主站蜘蛛池模板: 陕西省| 绥宁县| 高雄县| 神农架林区| 报价| 蓬安县| 青河县| 正阳县| 临湘市| 河南省| 思南县| 民丰县| 滦南县| 临城县| 建昌县| 馆陶县| 丹江口市| 常州市| 阿巴嘎旗| 繁峙县| 许昌市| 土默特右旗| 安达市| 驻马店市| 扎鲁特旗| 陕西省| 安远县| 咸宁市| 右玉县| 唐河县| 禄丰县| 通道| 乐东| 吴堡县| 绥棱县| 吉隆县| 霍州市| 阳泉市| 金沙县| 米林县| 梁山县|