Xử lý dữ liệu quan hệ với Hive trong Hadoop (07) - Julian M. Kunkel
Génération de l'aperçu...
Slide bài giảng về xử lý dữ liệu quan hệ với Hive trong Hadoop, bao gồm mô hình dữ liệu, thực thi truy vấn, định dạng file và HiveQL.
Description
Processing Relational Data with Hive Lecture BigData Analytics Julian M. Kunkel julian.kunkel@googlemail.com University of Hamburg / German Climate Computing Center (DKRZ) 2016-12-02 Disclaimer: Big Data software is constantly updated, code samples may be outdated. Hive: SQL in the Hadoop Environment Query Execution File Formats HiveQL Summary Outline 1 Hive: SQL in the Hadoop Environment 2 Query Execution 3 File Formats 4 HiveQL 5 Summary Julian M. Kunkel Lecture BigData Analytics, 2016 2 / 47 Hive: SQL in the Hadoop Environment Query Execution File Formats HiveQL Summary Hive Overview Hive: Data warehouse functionality on top of Hadoop/HDFS Compute engines: Map/Reduce, Tez, Spark Storage formats: Text, ORC, HBASE, RCFile, Avro Manages metadata (schemes) in RDBMS (or HBase) Access via: SQL-like query language HiveQL Similar to SQL-92 but several features are missing Limited transactions, subquery and views Query latency: 10s of seconds to minutes (new versions: sub-seconds) Features Basic data indexing (compaction and bitmaps) User-defined functions to manipulate data and support data-mining Interactive shell: hive Hive Web interface (simple GUI data access and query execution) WebHCat API (RESTful interface) Julian M. Kunkel Lecture BigData Analytics, 2016 3 / 47 Hive: SQL in the Hadoop Environment Query Execution File Formats HiveQL Summary Data Model [22] Data types Primitive types (int, float, strings, dates, boolean) Bags (arrays), dictionaries Derived data types (structs) can be defined by users Data organization Table: Like in relational databases with a schema The Hive data definition language (DDL) manages tables Data is stored in files on HDFS Partitions: table key determining the mapping to directories Reduces the amount of data to be accessed in filters Example key: /ds=<date> for table T Predicate T.ds=’2008-09-01’ searches for files in /ds=2008-09-01/ directory Buckets/Clusters: Data of partitions are mapped
Résumé IA
- Nom du document
- Xử lý dữ liệu quan hệ với Hive trong Hadoop (07) - Julian M. Kunkel
- École / Cours
- University of Hamburg · Big Data
- Auteur (dans le document)
- Julian M. Kunkel
- Contenu
- Tài liệu này cung cấp cái nhìn tổng quan về Hive, một công cụ kho dữ liệu trên Hadoop, tập trung vào cách truy vấn dữ liệu bằng HiveQL, mô hình dữ liệu, quản lý schema với HCatalog và các phương thức truy cập.
- Table des matières
- Hive: SQL in the Hadoop Environment
- Query Execution
- File Formats
- HiveQL
- Summary
- Pages
- 48 pages
- Téléversé par
- Uni24h
Foire aux questions
Ce document est-il gratuit ?
Oui. « Xử lý dữ liệu quan hệ với Hive trong Hadoop (07) - Julian M. Kunkel » est gratuit — il suffit de vous connecter et de cliquer sur Télécharger pour obtenir le fichier original.
Combien de pages compte ce document ?
Le document contient 48 pages, pour le cours Big Data. Vous pouvez le prévisualiser en ligne avant de le télécharger.
Puis-je prévisualiser avant de télécharger ?
Oui. Vous pouvez prévisualiser ce document directement sur cette page avec le lecteur en ligne, puis décider de le télécharger ou non.
Xử lý dữ liệu quan hệ với Hive trong Hadoop (07) - Julian M. Kunkel
Génération de l'aperçu...
Processing Relational Data with Hive Lecture BigData Analytics Julian M. Kunkel julian.kunkel@googlemail.com University of Hamburg / German Climate Computing Center (DKRZ) 2016-12-02 Disclaimer: Big Data software is constantly updated, code samples may be outdated. Hive: SQL in the Hadoop Environment Query Execution File Formats HiveQL Summary Outline 1 Hive: SQL in the Hadoop Environment 2 Query Execution 3 File Formats 4 HiveQL 5 Summary Julian M. Kunkel Lecture BigData Analytics, 2016 2 / 47 Hive: SQL in the Hadoop Environment Query Execution File Formats HiveQL Summary Hive Overview Hive: Data warehouse functionality on top of Hadoop/HDFS Compute engines: Map/Reduce, Tez, Spark Storage formats: Text, ORC, HBASE, RCFile, Avro Manages metadata (schemes) in RDBMS (or HBase) Access via: SQL-like query language HiveQL Similar to SQL-92 but several features are missing Limited transactions, subquery and views Query latency: 10s of seconds to minutes (new versions: sub-seconds) Features Basic data indexing (compaction and bitmaps) User-defined functions to manipulate data and support data-mining Interactive shell: hive Hive Web interface (simple GUI data access and query execution) WebHCat API (RESTful interface) Julian M. Kunkel Lecture BigData Analytics, 2016 3 / 47 Hive: SQL in the Hadoop Environment Query Execution File Formats HiveQL Summary Data Model [22] Data types Primitive types (int, float, strings, dates, boolean) Bags (arrays), dictionaries Derived data types (structs) can be defined by users Data organization Table: Like in relational databases with a schema The Hive data definition language (DDL) manages tables Data is stored in files on HDFS Partitions: table key determining the mapping to directories Reduces the amount of data to be accessed in filters Example key: /ds=<date> for table T Predicate T.ds=’2008-09-01’ searches for files in /ds=2008-09-01/ directory Buckets/Clusters: Data of partitions are mapped
Lire le document entier
- Nom du document
- Xử lý dữ liệu quan hệ với Hive trong Hadoop (07) - Julian M. Kunkel
- École / Cours
- University of Hamburg · Big Data
- Auteur (dans le document)
- Julian M. Kunkel
- Contenu
- Tài liệu này cung cấp cái nhìn tổng quan về Hive, một công cụ kho dữ liệu trên Hadoop, tập trung vào cách truy vấn dữ liệu bằng HiveQL, mô hình dữ liệu, quản lý schema với HCatalog và các phương thức truy cập.
- Table des matières
- Hive: SQL in the Hadoop Environment
- Query Execution
- File Formats
- HiveQL
- Summary
- Pages
- 48 pages
- Téléversé par
- Uni24h
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !