官术网_书友最值得收藏!

Data caching

Many machine learning algorithms are iterative in nature and thus require multiple passes over the data. However, all data stored in Spark RDD are by default transient, since RDD just stores the transformation to be executed and not the actual data. That means each action would recompute data again and again by executing the transformation stored in RDD.

Hence, Spark provides a way to persist the data in case we need to iterate over it. Spark also publishes several StorageLevels to allow storing data with various options:

  • NONE: No caching at all
  • MEMORY_ONLY: Caches RDD data only in memory
  • DISK_ONLY: Write cached RDD data to a disk and releases from memory
  • MEMORY_AND_DISK: Caches RDD in memory, if it's not possible to offload data to a disk
  • OFF_HEAP: Use external memory storage which is not part of JVM heap

Furthermore, Spark gives users the ability to cache data in two flavors: raw (for example, MEMORY_ONLY) and serialized (for example, MEMORY_ONLY_SER). The later uses large memory buffers to store serialized content of RDD directly. Which one to use is very task and resource dependent. A good rule of thumb is if the dataset you are working with is less than 10 gigs then raw caching is preferred to serialized caching. However, once you cross over the 10 gigs soft-threshold, raw caching imposes a greater memory footprint than serialized caching.

Spark can be forced to cache by calling the cache() method on RDD or directly via calling the method persist with the desired persistent target - persist(StorageLevels.MEMORY_ONLY_SER). It is useful to know that RDD allows us to set up the storage level only once.

The decision on what to cache and how to cache is part of the Spark magic; however, the golden rule is to use caching when we need to access RDD data several times and choose a destination based on the application preference respecting speed and storage. A great blogpost which goes into far more detail than what is given here is available at:

http://sujee.net/2015/01/22/understanding-spark-caching/#.VpU1nJMrLdc

Cached RDDs can be accessed as well from the H2O Flow UI by evaluating the cell with getRDDs:

主站蜘蛛池模板: 鹿泉市| 刚察县| 龙游县| 湖南省| 合肥市| 五寨县| 鄂托克前旗| 清水河县| 太白县| 安乡县| 喀喇沁旗| 菏泽市| 镇宁| 涡阳县| 泸西县| 宝坻区| 田阳县| 隆林| 巴彦淖尔市| 志丹县| 黄浦区| 稻城县| 丹凤县| 天门市| 岳阳市| 堆龙德庆县| 凌源市| 瑞金市| 宁德市| 全州县| 鄄城县| 赤峰市| 交城县| 九龙县| 东源县| 安宁市| 光泽县| 佛坪县| 龙江县| 长顺县| 金山区|