官术网_书友最值得收藏!

Data caching

Many machine learning algorithms are iterative in nature and thus require multiple passes over the data. However, all data stored in Spark RDD are by default transient, since RDD just stores the transformation to be executed and not the actual data. That means each action would recompute data again and again by executing the transformation stored in RDD.

Hence, Spark provides a way to persist the data in case we need to iterate over it. Spark also publishes several StorageLevels to allow storing data with various options:

  • NONE: No caching at all
  • MEMORY_ONLY: Caches RDD data only in memory
  • DISK_ONLY: Write cached RDD data to a disk and releases from memory
  • MEMORY_AND_DISK: Caches RDD in memory, if it's not possible to offload data to a disk
  • OFF_HEAP: Use external memory storage which is not part of JVM heap

Furthermore, Spark gives users the ability to cache data in two flavors: raw (for example, MEMORY_ONLY) and serialized (for example, MEMORY_ONLY_SER). The later uses large memory buffers to store serialized content of RDD directly. Which one to use is very task and resource dependent. A good rule of thumb is if the dataset you are working with is less than 10 gigs then raw caching is preferred to serialized caching. However, once you cross over the 10 gigs soft-threshold, raw caching imposes a greater memory footprint than serialized caching.

Spark can be forced to cache by calling the cache() method on RDD or directly via calling the method persist with the desired persistent target - persist(StorageLevels.MEMORY_ONLY_SER). It is useful to know that RDD allows us to set up the storage level only once.

The decision on what to cache and how to cache is part of the Spark magic; however, the golden rule is to use caching when we need to access RDD data several times and choose a destination based on the application preference respecting speed and storage. A great blogpost which goes into far more detail than what is given here is available at:

http://sujee.net/2015/01/22/understanding-spark-caching/#.VpU1nJMrLdc

Cached RDDs can be accessed as well from the H2O Flow UI by evaluating the cell with getRDDs:

主站蜘蛛池模板: 当涂县| 衢州市| 惠东县| 香港| 江永县| 通辽市| 钟山县| 姚安县| 萍乡市| 新田县| 宁陕县| 普格县| 田阳县| 威远县| 东乌| 定西市| 安泽县| 射阳县| 含山县| 舒兰市| 万安县| 廉江市| 枣阳市| 如东县| 汽车| 长子县| 武邑县| 呈贡县| 驻马店市| 婺源县| 甘泉县| 潮州市| 卢湾区| 珲春市| 密山市| 错那县| 泽州县| 卢氏县| 广宗县| 宁波市| 瑞金市|