官术网_书友最值得收藏!

Optimizing the Storing and Processing of Data for Machine Learning Problems

All of the preceding uses for artificial intelligence rely heavily on optimized data storage and processing. Optimization is necessary for machine learning because the data size can be huge, as seen in the following examples:

  • A single X-ray file can be many gigabytes in size.
  • Translation corpora (large collections of texts) can reach billions of sentences.
  • YouTube's stored data is measured in exabytes.
  • Financial data might seem like just a few numbers; these are generated in such large quantities per second that the New York Stock Exchange generates 1 TB of data daily.

While every machine learning system is unique, in many systems, data touches the same components. In a hypothetical machine learning system, data might be dealt with as follows:

Figure 1.4: Hardware used in a hypothetical machine learning system

Each of these is a highly specialized piece of hardware, and although not all of them store data for long periods in the way traditional hard disks or tape backups do, it is important to know how data storage can be optimized at each stage. Let's pe into a text classification AI project to see how optimizations can be applied at some stages.

主站蜘蛛池模板: 麻阳| 宁晋县| 克什克腾旗| 天峻县| 徐水县| 那曲县| 涿鹿县| 安宁市| 东平县| 天镇县| 清新县| 扎赉特旗| 亚东县| 突泉县| 南雄市| 澄迈县| 隆尧县| 图木舒克市| 启东市| 长沙县| 平谷区| 盐源县| 定边县| 堆龙德庆县| 静乐县| 长宁区| 吉林市| 东丰县| 马鞍山市| 太康县| 东方市| 远安县| 寿光市| 池州市| 宁化县| 凤城市| 襄樊市| 连城县| 元江| 琼结县| 无为县|