官术网_书友最值得收藏!

The basics of web requests

The worldwide capacity to generate data is estimated to double in size every two years. Even though there is an interdisciplinary field known as data science that is entirely dedicated to the study of data, almost every programming task in software development also has something to do with collecting and analyzing data. A significant part of this is, of course, data collection. However, the data that we need for our applications is sometimes not stored nicely and cleanly in a database—sometimes, we need to collect the data we need from web pages.

For example, web scraping is a data extraction method that automatically makes requests to web pages and downloads specific information. Web scraping allows us to comb through numerous websites and collect any data we need in a systematic and consistent manner—the collected data can be analyzed later on by our applications or simply saved on our computers in various formats. An example of this would be Google, which programs and runs numerous web scrapers of its own to find and index web pages for the search engine.

The Python language itself provides a number of good options for applications of this kind. In this chapter, we will mainly work with the requests module to make client-side web requests from our Python programs. However, before we look into this module in more detail, we need to understand some web terminology in order to be able to effectively design our applications.

主站蜘蛛池模板: 什邡市| 高安市| 波密县| 赤城县| 新巴尔虎右旗| 灵宝市| 桂东县| 赣榆县| 海宁市| 邳州市| 阿巴嘎旗| 曲周县| 朔州市| 平乡县| 东源县| 西林县| 抚顺县| 花垣县| 刚察县| 蒙阴县| 卓尼县| 安平县| 云梦县| 越西县| 北川| 天长市| 富宁县| 怀宁县| 政和县| 延边| 苗栗市| 大安市| 南通市| 桃源县| 泗洪县| 上饶市| 巢湖市| 扎赉特旗| 横山县| 海口市| 电白县|