Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
Génération de l'aperçu...
Slide bài giảng về tối ưu hóa và thực thi của Pig, giới thiệu kiến trúc và các kỹ thuật tối ưu như giảm quét, giảm số lượng MR job, giảm shuffle, xử lý skew, tối ưu bộ nhớ, và các mô hình thực thi cải tiến.
Description
Pig Optimization and Execution Alan F. Gates @alanfgates © Hortonworks Inc. 2011 Page 1 Who Am I? Pig committer and PMC Member HCatalog committer and mentor Member of ASF and Incubator PMC Co-founder of Hortonworks Author of Programming Pig from O’Reilly Photo credit: Steven Guarnaccia, The Three Little Pigs Who Are You? 3 What Should We Optimize? Minimize scans – Hadoop is still often I/O bound Minimize total number of MR jobs Minimize shuffle size and number of shuffles Avoid spills to disk Reduce or remove skew For small jobs, minimize start-up time 4 Pig Deployment No server, all optimization and planning done on the launching machine Job executes on cluster Pig resides on user machine or gateway Hadoop Cluster User machine Pig Guts (i.e. Pig Architecture), p. 1 Logical Plan Pig Latin Load A = LOAD ‘myfile’ AS (x, y, z); B = GROUP A by x; C = FILTER B by group > 0; D = FOREACH C GENERATE group, COUNT(A); STORE D INTO ‘output’; Group AST Filter Foreach Semantic Checks Store 6 Pig Guts, p. 2 MapReduce Plan Logical Plan Load Load Group Filter Filter Foreach Map Filter Rearrange Group Rule based optimizations Reduce Store 7 Foreach Package Store Foreach Pig Guts, p. 3 MapReduce Plan Map Map Filter Filter Rearrange Rearrange Combine Foreach Physical optimizations Reduce Package Reduce Package Foreach Foreach 8 It would be really cool if… Map Map Reduce Reduce Map Map Reduce Reduce Map Reduce What’s the right join algorithm here? Even with statistics it would be hard to know. Need on the fly execution plan rewrites. 9 Memory Java + memory management = oil + water Java types inefficient memory users (~4x disk size) Very difficult to tell how much memory you are using Originally tried to monitor memory use via MXBeans: FAIL! Now estimate number of records we can hold in memory and spill when we exceed; allow user to tune guess 10 Reducing Spills to Disk Select Map size an
Résumé IA
- Nom du document
- Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
- École / Cours
- Duke University · Big Data
- Auteur (dans le document)
- Alan F. Gates
- Contenu
- Tài liệu này cung cấp cái nhìn sâu sắc về kiến trúc và các phương pháp tối ưu hóa hiệu suất cho Apache Pig, bao gồm giảm thiểu I/O, số lượng job MapReduce, và xử lý dữ liệu lệch. Nó cũng thảo luận về quản lý bộ nhớ, khởi động job nhanh hơn và các mô hình thực thi cải tiến.
- Table des matières
- Who Am I?
- Who Are You?
- What Should We Optimize?
- Pig Deployment
- Pig Guts (i.e. Pig Architecture), p. 1
- Pig Guts, p. 2
- Pig Guts, p. 3
- It would be really cool if…
- Memory
- Reducing Spills to Disk
- Skew
- Reducing your Reducers
- (De)serialization
- Faster Job Startup
- Improved Execution Models
- Code Generation
- Learn More
- Questions?
- Pages
- 19 pages
- Téléversé par
- Uni24h
Foire aux questions
Ce document est-il gratuit ?
Oui. « Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh) » est gratuit — il suffit de vous connecter et de cliquer sur Télécharger pour obtenir le fichier original.
Combien de pages compte ce document ?
Le document contient 19 pages, pour le cours Big Data. Vous pouvez le prévisualiser en ligne avant de le télécharger.
Puis-je prévisualiser avant de télécharger ?
Oui. Vous pouvez prévisualiser ce document directement sur cette page avec le lecteur en ligne, puis décider de le télécharger ou non.
Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
Génération de l'aperçu...
Pig Optimization and Execution Alan F. Gates @alanfgates © Hortonworks Inc. 2011 Page 1 Who Am I? Pig committer and PMC Member HCatalog committer and mentor Member of ASF and Incubator PMC Co-founder of Hortonworks Author of Programming Pig from O’Reilly Photo credit: Steven Guarnaccia, The Three Little Pigs Who Are You? 3 What Should We Optimize? Minimize scans – Hadoop is still often I/O bound Minimize total number of MR jobs Minimize shuffle size and number of shuffles Avoid spills to disk Reduce or remove skew For small jobs, minimize start-up time 4 Pig Deployment No server, all optimization and planning done on the launching machine Job executes on cluster Pig resides on user machine or gateway Hadoop Cluster User machine Pig Guts (i.e. Pig Architecture), p. 1 Logical Plan Pig Latin Load A = LOAD ‘myfile’ AS (x, y, z); B = GROUP A by x; C = FILTER B by group > 0; D = FOREACH C GENERATE group, COUNT(A); STORE D INTO ‘output’; Group AST Filter Foreach Semantic Checks Store 6 Pig Guts, p. 2 MapReduce Plan Logical Plan Load Load Group Filter Filter Foreach Map Filter Rearrange Group Rule based optimizations Reduce Store 7 Foreach Package Store Foreach Pig Guts, p. 3 MapReduce Plan Map Map Filter Filter Rearrange Rearrange Combine Foreach Physical optimizations Reduce Package Reduce Package Foreach Foreach 8 It would be really cool if… Map Map Reduce Reduce Map Map Reduce Reduce Map Reduce What’s the right join algorithm here? Even with statistics it would be hard to know. Need on the fly execution plan rewrites. 9 Memory Java + memory management = oil + water Java types inefficient memory users (~4x disk size) Very difficult to tell how much memory you are using Originally tried to monitor memory use via MXBeans: FAIL! Now estimate number of records we can hold in memory and spill when we exceed; allow user to tune guess 10 Reducing Spills to Disk Select Map size an
Lire le document entier
- Nom du document
- Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
- École / Cours
- Duke University · Big Data
- Auteur (dans le document)
- Alan F. Gates
- Contenu
- Tài liệu này cung cấp cái nhìn sâu sắc về kiến trúc và các phương pháp tối ưu hóa hiệu suất cho Apache Pig, bao gồm giảm thiểu I/O, số lượng job MapReduce, và xử lý dữ liệu lệch. Nó cũng thảo luận về quản lý bộ nhớ, khởi động job nhanh hơn và các mô hình thực thi cải tiến.
- Table des matières
- Who Am I?
- Who Are You?
- What Should We Optimize?
- Pig Deployment
- Pig Guts (i.e. Pig Architecture), p. 1
- Pig Guts, p. 2
- Pig Guts, p. 3
- It would be really cool if…
- Memory
- Reducing Spills to Disk
- Skew
- Reducing your Reducers
- (De)serialization
- Faster Job Startup
- Improved Execution Models
- Code Generation
- Learn More
- Questions?
- Pages
- 19 pages
- Téléversé par
- Uni24h
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !