Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
Generating preview...
Slide bài giảng về tối ưu hóa và thực thi của Pig, giới thiệu kiến trúc và các kỹ thuật tối ưu như giảm quét, giảm số lượng MR job, giảm shuffle, xử lý skew, tối ưu bộ nhớ, và các mô hình thực thi cải tiến.
Description
Pig Optimization and Execution Alan F. Gates @alanfgates © Hortonworks Inc. 2011 Page 1 Who Am I? Pig committer and PMC Member HCatalog committer and mentor Member of ASF and Incubator PMC Co-founder of Hortonworks Author of Programming Pig from O’Reilly Photo credit: Steven Guarnaccia, The Three Little Pigs Who Are You? 3 What Should We Optimize? Minimize scans – Hadoop is still often I/O bound Minimize total number of MR jobs Minimize shuffle size and number of shuffles Avoid spills to disk Reduce or remove skew For small jobs, minimize start-up time 4 Pig Deployment No server, all optimization and planning done on the launching machine Job executes on cluster Pig resides on user machine or gateway Hadoop Cluster User machine Pig Guts (i.e. Pig Architecture), p. 1 Logical Plan Pig Latin Load A = LOAD ‘myfile’ AS (x, y, z); B = GROUP A by x; C = FILTER B by group > 0; D = FOREACH C GENERATE group, COUNT(A); STORE D INTO ‘output’; Group AST Filter Foreach Semantic Checks Store 6 Pig Guts, p. 2 MapReduce Plan Logical Plan Load Load Group Filter Filter Foreach Map Filter Rearrange Group Rule based optimizations Reduce Store 7 Foreach Package Store Foreach Pig Guts, p. 3 MapReduce Plan Map Map Filter Filter Rearrange Rearrange Combine Foreach Physical optimizations Reduce Package Reduce Package Foreach Foreach 8 It would be really cool if… Map Map Reduce Reduce Map Map Reduce Reduce Map Reduce What’s the right join algorithm here? Even with statistics it would be hard to know. Need on the fly execution plan rewrites. 9 Memory Java + memory management = oil + water Java types inefficient memory users (~4x disk size) Very difficult to tell how much memory you are using Originally tried to monitor memory use via MXBeans: FAIL! Now estimate number of records we can hold in memory and spill when we exceed; allow user to tune guess 10 Reducing Spills to Disk Select Map size an
AI summary
- Document name
- Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
- School / Course
- Duke University · Big Data
- Author (in document)
- Alan F. Gates
- Content
- Tài liệu này cung cấp cái nhìn sâu sắc về kiến trúc và các phương pháp tối ưu hóa hiệu suất cho Apache Pig, bao gồm giảm thiểu I/O, số lượng job MapReduce, và xử lý dữ liệu lệch. Nó cũng thảo luận về quản lý bộ nhớ, khởi động job nhanh hơn và các mô hình thực thi cải tiến.
- Table of contents
- Who Am I?
- Who Are You?
- What Should We Optimize?
- Pig Deployment
- Pig Guts (i.e. Pig Architecture), p. 1
- Pig Guts, p. 2
- Pig Guts, p. 3
- It would be really cool if…
- Memory
- Reducing Spills to Disk
- Skew
- Reducing your Reducers
- (De)serialization
- Faster Job Startup
- Improved Execution Models
- Code Generation
- Learn More
- Questions?
- Pages
- 19 pages
- Uploaded by
- Uni24h
Frequently asked questions
Is this document free?
Yes. “Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)” is free — just sign in and click Download to get the original file.
How many pages is this document?
The document has 19 pages, for the course Big Data. You can preview it online before downloading.
Can I preview before downloading?
Yes. You can preview this document right on this page with the online reader, then decide whether to download.
Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
Generating preview...
Pig Optimization and Execution Alan F. Gates @alanfgates © Hortonworks Inc. 2011 Page 1 Who Am I? Pig committer and PMC Member HCatalog committer and mentor Member of ASF and Incubator PMC Co-founder of Hortonworks Author of Programming Pig from O’Reilly Photo credit: Steven Guarnaccia, The Three Little Pigs Who Are You? 3 What Should We Optimize? Minimize scans – Hadoop is still often I/O bound Minimize total number of MR jobs Minimize shuffle size and number of shuffles Avoid spills to disk Reduce or remove skew For small jobs, minimize start-up time 4 Pig Deployment No server, all optimization and planning done on the launching machine Job executes on cluster Pig resides on user machine or gateway Hadoop Cluster User machine Pig Guts (i.e. Pig Architecture), p. 1 Logical Plan Pig Latin Load A = LOAD ‘myfile’ AS (x, y, z); B = GROUP A by x; C = FILTER B by group > 0; D = FOREACH C GENERATE group, COUNT(A); STORE D INTO ‘output’; Group AST Filter Foreach Semantic Checks Store 6 Pig Guts, p. 2 MapReduce Plan Logical Plan Load Load Group Filter Filter Foreach Map Filter Rearrange Group Rule based optimizations Reduce Store 7 Foreach Package Store Foreach Pig Guts, p. 3 MapReduce Plan Map Map Filter Filter Rearrange Rearrange Combine Foreach Physical optimizations Reduce Package Reduce Package Foreach Foreach 8 It would be really cool if… Map Map Reduce Reduce Map Map Reduce Reduce Map Reduce What’s the right join algorithm here? Even with statistics it would be hard to know. Need on the fly execution plan rewrites. 9 Memory Java + memory management = oil + water Java types inefficient memory users (~4x disk size) Very difficult to tell how much memory you are using Originally tried to monitor memory use via MXBeans: FAIL! Now estimate number of records we can hold in memory and spill when we exceed; allow user to tune guess 10 Reducing Spills to Disk Select Map size an
Read full document
- Document name
- Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
- School / Course
- Duke University · Big Data
- Author (in document)
- Alan F. Gates
- Content
- Tài liệu này cung cấp cái nhìn sâu sắc về kiến trúc và các phương pháp tối ưu hóa hiệu suất cho Apache Pig, bao gồm giảm thiểu I/O, số lượng job MapReduce, và xử lý dữ liệu lệch. Nó cũng thảo luận về quản lý bộ nhớ, khởi động job nhanh hơn và các mô hình thực thi cải tiến.
- Table of contents
- Who Am I?
- Who Are You?
- What Should We Optimize?
- Pig Deployment
- Pig Guts (i.e. Pig Architecture), p. 1
- Pig Guts, p. 2
- Pig Guts, p. 3
- It would be really cool if…
- Memory
- Reducing Spills to Disk
- Skew
- Reducing your Reducers
- (De)serialization
- Faster Job Startup
- Improved Execution Models
- Code Generation
- Learn More
- Questions?
- Pages
- 19 pages
- Uploaded by
- Uni24h
Comments (0)
No comments yet. Be the first!
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Comments (0)
No comments yet. Be the first!