官术网_书友最值得收藏!

Optimizing the Storing and Processing of Data for Machine Learning Problems

All of the preceding uses for artificial intelligence rely heavily on optimized data storage and processing. Optimization is necessary for machine learning because the data size can be huge, as seen in the following examples:

  • A single X-ray file can be many gigabytes in size.
  • Translation corpora (large collections of texts) can reach billions of sentences.
  • YouTube's stored data is measured in exabytes.
  • Financial data might seem like just a few numbers; these are generated in such large quantities per second that the New York Stock Exchange generates 1 TB of data daily.

While every machine learning system is unique, in many systems, data touches the same components. In a hypothetical machine learning system, data might be dealt with as follows:

Figure 1.4: Hardware used in a hypothetical machine learning system

Each of these is a highly specialized piece of hardware, and although not all of them store data for long periods in the way traditional hard disks or tape backups do, it is important to know how data storage can be optimized at each stage. Let's pe into a text classification AI project to see how optimizations can be applied at some stages.

主站蜘蛛池模板: 江油市| 马公市| 肇庆市| 青冈县| 正镶白旗| 武冈市| 金堂县| 容城县| 邹平县| 潼南县| 邵东县| 孝感市| 巢湖市| 连山| 宁晋县| 锡林郭勒盟| 开原市| 宁城县| 修武县| 盐亭县| 洛扎县| 稷山县| 西贡区| 肇庆市| 天门市| 清远市| 蛟河市| 镇赉县| 德惠市| 金湖县| 洛阳市| 泗水县| 浮山县| 山阳县| 溆浦县| 多伦县| 云安县| 股票| 阿拉善右旗| 新绛县| 仁布县|