qp joins (07) (Thực thi truy vấn trong hệ thống tính toán dữ liệu chuyên sâu) (Tiếng Anh)
Génération de l'aperçu...
Slide bài giảng về thực thi truy vấn trong các hệ thống tính toán dữ liệu chuyên sâu, tập trung vào các toán tử sắp xếp và kết nối.
Description
Data-Intensive Computing Systems Query Execution (Sort and Join operators) Shivnath Babu Roadmap A simple operator: Nested Loop Join Preliminaries Cost model Clustering Operator classes Operator implementation (with examples from joins) Scan-based Sort-based Using existing indexes Hash-based Buffer Management Parallel Processing Nested Loop Join (NLJ) R1 B a a b d C 10 20 10 30 C 10 40 15 20 D cat dog bat rat NLJ (conceptually) for each r R1 do for each s R2 do if r.C = s.C then output r,s pair R2 Nested Loop Join (contd.) Tuple-based Block-based Asymmetric Implementing Operators Basic algorithm Scan-based (e.g., NLJ) Sort-based Using existing indexes Hash-based (building an index on the fly) Memory management Tradeoff between memory and #IOs Parallel processing Roadmap A simple operator: Nested Loop Join Preliminaries Cost model Clustering Operator classes Operator implementation (with examples from joins) Scan-based Sort-based Using existing indexes Hash-based Buffer Management Parallel Processing Operator Cost Model Simplest: Count # of disk blocks read and written during operator execution Extends to query plans Cost of query plan = Sum of operator costs Caution: Ignoring CPU costs Assumptions Single-processor-single-disk machine Will consider parallelism later Ignore cost of writing out result Output size is independent of operator implementation Ignore # accesses to index blocks Parameters used in Cost Model B(R) = # blocks storing R tuples T(R) = # tuples in R V(R,A) = # distinct values of attr A in R M = # memory blocks available Roadmap A simple operator: Nested Loop Join Preliminaries Cost model Clustering Operator classes Operator implementation (with examples from joins) Scan-based Sort-based Using existing indexes Hash-based Buffer Management Parallel Processing Notions of clustering Clustered file o
Résumé IA
- Nom du document
- qp joins (07) (Thực thi truy vấn trong hệ thống tính toán dữ liệu chuyên sâu) (Tiếng Anh)
- École / Cours
- Duke University · Big Data
- Contenu
- Tài liệu giới thiệu các toán tử thực thi truy vấn, bắt đầu với Nested Loop Join, sau đó thảo luận về mô hình chi phí, phân cụm, và các phương pháp triển khai toán tử khác nhau như dựa trên quét, sắp xếp, chỉ mục và băm.
- Table des matières
- Roadmap
- Nested Loop Join (NLJ)
- Nested Loop Join (contd.)
- Implementing Operators
- Roadmap
- Operator Cost Model
- Assumptions
- Parameters used in Cost Model
- Roadmap
- Notions of clustering
- Clustering Index
- Examples
- Operator Classes
- Roadmap
- Implementing Tuple-at-a-time Operators
- Implementing a Full-Relation Operator, Ex: Sort
- Implementing a Full-Relation Operator, Ex: Sort
- Two-phase Sort: Phase 1
- Two-phase Sort: Phase 2
- Analysis of Two-Phase Sort
- Pages
- 77 pages
- Téléversé par
- Uni24h
Foire aux questions
Ce document est-il gratuit ?
Oui. « qp joins (07) (Thực thi truy vấn trong hệ thống tính toán dữ liệu chuyên sâu) (Tiếng Anh) » est gratuit — il suffit de vous connecter et de cliquer sur Télécharger pour obtenir le fichier original.
Combien de pages compte ce document ?
Le document contient 77 pages, pour le cours Big Data. Vous pouvez le prévisualiser en ligne avant de le télécharger.
Puis-je prévisualiser avant de télécharger ?
Oui. Vous pouvez prévisualiser ce document directement sur cette page avec le lecteur en ligne, puis décider de le télécharger ou non.
qp joins (07) (Thực thi truy vấn trong hệ thống tính toán dữ liệu chuyên sâu) (Tiếng Anh)
Génération de l'aperçu...
Data-Intensive Computing Systems Query Execution (Sort and Join operators) Shivnath Babu Roadmap A simple operator: Nested Loop Join Preliminaries Cost model Clustering Operator classes Operator implementation (with examples from joins) Scan-based Sort-based Using existing indexes Hash-based Buffer Management Parallel Processing Nested Loop Join (NLJ) R1 B a a b d C 10 20 10 30 C 10 40 15 20 D cat dog bat rat NLJ (conceptually) for each r R1 do for each s R2 do if r.C = s.C then output r,s pair R2 Nested Loop Join (contd.) Tuple-based Block-based Asymmetric Implementing Operators Basic algorithm Scan-based (e.g., NLJ) Sort-based Using existing indexes Hash-based (building an index on the fly) Memory management Tradeoff between memory and #IOs Parallel processing Roadmap A simple operator: Nested Loop Join Preliminaries Cost model Clustering Operator classes Operator implementation (with examples from joins) Scan-based Sort-based Using existing indexes Hash-based Buffer Management Parallel Processing Operator Cost Model Simplest: Count # of disk blocks read and written during operator execution Extends to query plans Cost of query plan = Sum of operator costs Caution: Ignoring CPU costs Assumptions Single-processor-single-disk machine Will consider parallelism later Ignore cost of writing out result Output size is independent of operator implementation Ignore # accesses to index blocks Parameters used in Cost Model B(R) = # blocks storing R tuples T(R) = # tuples in R V(R,A) = # distinct values of attr A in R M = # memory blocks available Roadmap A simple operator: Nested Loop Join Preliminaries Cost model Clustering Operator classes Operator implementation (with examples from joins) Scan-based Sort-based Using existing indexes Hash-based Buffer Management Parallel Processing Notions of clustering Clustered file o
Lire le document entier
- Nom du document
- qp joins (07) (Thực thi truy vấn trong hệ thống tính toán dữ liệu chuyên sâu) (Tiếng Anh)
- École / Cours
- Duke University · Big Data
- Contenu
- Tài liệu giới thiệu các toán tử thực thi truy vấn, bắt đầu với Nested Loop Join, sau đó thảo luận về mô hình chi phí, phân cụm, và các phương pháp triển khai toán tử khác nhau như dựa trên quét, sắp xếp, chỉ mục và băm.
- Table des matières
- Roadmap
- Nested Loop Join (NLJ)
- Nested Loop Join (contd.)
- Implementing Operators
- Roadmap
- Operator Cost Model
- Assumptions
- Parameters used in Cost Model
- Roadmap
- Notions of clustering
- Clustering Index
- Examples
- Operator Classes
- Roadmap
- Implementing Tuple-at-a-time Operators
- Implementing a Full-Relation Operator, Ex: Sort
- Implementing a Full-Relation Operator, Ex: Sort
- Two-phase Sort: Phase 1
- Two-phase Sort: Phase 2
- Analysis of Two-Phase Sort
- Pages
- 77 pages
- Téléversé par
- Uni24h
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !