官术网_书友最值得收藏!

Data caching

Many machine learning algorithms are iterative in nature and thus require multiple passes over the data. However, all data stored in Spark RDD are by default transient, since RDD just stores the transformation to be executed and not the actual data. That means each action would recompute data again and again by executing the transformation stored in RDD.

Hence, Spark provides a way to persist the data in case we need to iterate over it. Spark also publishes several StorageLevels to allow storing data with various options:

  • NONE: No caching at all
  • MEMORY_ONLY: Caches RDD data only in memory
  • DISK_ONLY: Write cached RDD data to a disk and releases from memory
  • MEMORY_AND_DISK: Caches RDD in memory, if it's not possible to offload data to a disk
  • OFF_HEAP: Use external memory storage which is not part of JVM heap

Furthermore, Spark gives users the ability to cache data in two flavors: raw (for example, MEMORY_ONLY) and serialized (for example, MEMORY_ONLY_SER). The later uses large memory buffers to store serialized content of RDD directly. Which one to use is very task and resource dependent. A good rule of thumb is if the dataset you are working with is less than 10 gigs then raw caching is preferred to serialized caching. However, once you cross over the 10 gigs soft-threshold, raw caching imposes a greater memory footprint than serialized caching.

Spark can be forced to cache by calling the cache() method on RDD or directly via calling the method persist with the desired persistent target - persist(StorageLevels.MEMORY_ONLY_SER). It is useful to know that RDD allows us to set up the storage level only once.

The decision on what to cache and how to cache is part of the Spark magic; however, the golden rule is to use caching when we need to access RDD data several times and choose a destination based on the application preference respecting speed and storage. A great blogpost which goes into far more detail than what is given here is available at:

http://sujee.net/2015/01/22/understanding-spark-caching/#.VpU1nJMrLdc

Cached RDDs can be accessed as well from the H2O Flow UI by evaluating the cell with getRDDs:

主站蜘蛛池模板: 雷山县| 攀枝花市| 潞西市| 额尔古纳市| 勐海县| 开阳县| 汉川市| 辽阳市| 弥渡县| 沛县| 尚志市| 苗栗市| 江川县| 永修县| 新津县| 巴林右旗| 长顺县| 中卫市| 永春县| 郁南县| 额敏县| 沈阳市| 海口市| 永仁县| 县级市| 商洛市| 洛川县| 错那县| 和静县| 潍坊市| 如东县| 肇东市| 镇宁| 大埔县| 仁布县| 弋阳县| 象山县| 砚山县| 罗田县| 沽源县| 壶关县|