官术网_书友最值得收藏!

Datasets and modeling

We're going to be using two of the prior datasets, the simulated data from Chapter 4Advanced Feature Selection in Linear Models, and the customer satisfaction data from Chapter 3, Logistic Regression. We'll start by building a classification tree on the simulated data. This will help us to understand the basic principles of tree-based methods. Then, we'll move on to random forest and boosted trees applied to the customer satisfaction data. This exercise will provide an excellent comparison to the generalized linear models from before. Finally, I want to show you an interesting feature selection method using random forest, using the simulated data. By interesting, I mean it's a valuable technique to add to your feature selection arsenal, but I'll point out a couple of caveats for you to consider in practical application.

主站蜘蛛池模板: 泰来县| 杨浦区| 雷山县| 金溪县| 金沙县| 绥阳县| 同仁县| 千阳县| 乌鲁木齐市| 泰安市| 莱州市| 玉环县| 景德镇市| 瑞昌市| 屏南县| 崇礼县| 洱源县| 金秀| 双柏县| 琼海市| 东乌珠穆沁旗| 浦东新区| 巴马| 彭泽县| 吉木萨尔县| 酒泉市| 资中县| 涿鹿县| 桦甸市| 河源市| 万源市| 临汾市| 濉溪县| 重庆市| 天镇县| 阿荣旗| 沾化县| 汕头市| 通渭县| 高雄市| 云林县|