書名： Mastering Apache Spark 2.x（Second Edition）
作者名： Romeo Kienzler
本章字?jǐn)?shù)： 127字
更新時(shí)間： 2021-07-02 18:55:32

Managing temporary views with the catalog API

Since Apache Spark 2.0, the catalog API is used to create and remove temporary views from an internal meta store. This is necessary if you want to use SQL, because it basically provides the mapping between a virtual table name and a DataFrame or Dataset.

Internally, Apache Spark uses the org.apache.spark.sql.catalyst.catalog.SessionCatalog class to manage temporary views as well as persistent tables.

Temporary views are stored in the SparkSession object, as persistent tables are stored in an external metastore. The abstract base class org.apache.spark.sql.catalyst.catalog.ExternalCatalog is extended for various meta store providers. One already exists for using Apache Derby and another one for the Apache Hive metastore, but anyone could extend this class and make Apache Spark use another metastore as well.

官术网_书友最值得收藏!

Mastering Apache Spark 2.x（Second Edition）

Managing temporary views with the catalog API