Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
- Seiten
- 19
- Định dạng
- PPTX
- Dung lượng
- 433 KB
- Năm
- 2011
- Trường
- Duke University
- Aufrufe
- 0
- Kommentare
- 0
- Lượt tải
- 0
Vorschau wird generiert...
Slide bài giảng về tối ưu hóa và thực thi của Pig, giới thiệu kiến trúc và các kỹ thuật tối ưu như giảm quét, giảm số lượng MR job, giảm shuffle, xử lý skew, tối ưu bộ nhớ, và các mô hình thực thi cải tiến.
- Dokumentenname
- Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
- Schule / Kurs
- Duke University · Big Data
- Autor (im Dokument)
- Alan F. Gates
- Inhalt
- Tài liệu này cung cấp cái nhìn sâu sắc về kiến trúc và các phương pháp tối ưu hóa hiệu suất cho Apache Pig, bao gồm giảm thiểu I/O, số lượng job MapReduce, và xử lý dữ liệu lệch. Nó cũng thảo luận về quản lý bộ nhớ, khởi động job nhanh hơn và các mô hình thực thi cải tiến.
- Inhaltsverzeichnis
- Who Am I?
- Who Are You?
- What Should We Optimize?
- Pig Deployment
- Pig Guts (i.e. Pig Architecture), p. 1
- Pig Guts, p. 2
- Pig Guts, p. 3
- It would be really cool if…
- Memory
- Reducing Spills to Disk
- Skew
- Reducing your Reducers
- (De)serialization
- Faster Job Startup
- Improved Execution Models
- Code Generation
- Learn More
- Questions?
- Seiten
- 19 Seiten
- Hochgeladen von
- Uni24h
Beschreibung
Trích nội dung tài liệu
Pig Optimization and Execution Alan F. Gates @alanfgates © Hortonworks Inc. 2011 Page 1 Who Am I? Pig committer and PMC Member HCatalog committer and mentor Member of ASF and Incubator PMC Co-founder of Hortonworks Author of Programming Pig from O’Reilly Photo credit: Steven Guarnaccia, The Three Little Pigs Who Are You? 3 What Should We Optimize? Minimize scans – Hadoop is still often I/O bound Minimize total number of MR jobs Minimize shuffle size and number of shuffles Avoid spills to disk Reduce or remove skew For small jobs, minimize start-up time 4 Pig Deployment No server, all optimization and planning done on the launching machine Job executes on cluster Pig resides on user machine or gateway Hadoop Cluster User machine Pig Guts (i.e. Pig Architecture), p. 1 Logical Plan Pig Latin Load A = LOAD ‘myfile’ AS (x, y, z); B = GROUP A by x; C = FILTER B by group > 0; D = FOREACH C GENERATE group, COUNT(A); STORE D INTO ‘output’; Group AST Filter Foreach Semantic Checks Store 6 Pig Guts, p. 2 MapReduce Plan Logical Plan Load Load Group Filter Filter Foreach Map Filter Rearrange Group Rule based optimizations Reduce Store 7 Foreach Package Store Foreach Pig Guts, p. 3 MapReduce Plan Map Map Filter Filter Rearrange Rearrange Combine Foreach Physical optimizations Reduce Package Reduce Package Foreach Foreach 8 It would be really cool if… Map Map Reduce Reduce Map Map Reduce Reduce Map Reduce What’s the right join algorithm here? Even with statistics it would be hard to know. Need on the fly execution plan rewrites. 9 Memory Java + memory management = oil + water Java types inefficient memory users (~4x disk size) Very difficult to tell how much memory you are using Originally tried to monitor memory use via MXBeans: FAIL! Now estimate number of records we can hold in memory and spill when we exceed; allow user to tune guess 10 Reducing Spills to Disk Select Map size an
Häufig gestellte Fragen
Ist dieses Dokument kostenlos?
Ja. „Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)“ ist kostenlos — melden Sie sich einfach an und klicken Sie auf Herunterladen, um die Originaldatei zu erhalten.
Wie viele Seiten hat dieses Dokument?
Das Dokument hat 19 Seiten, für den Kurs Big Data. Sie können es vor dem Herunterladen online in der Vorschau ansehen.
Kann ich vor dem Herunterladen eine Vorschau ansehen?
Ja. Sie können sich dieses Dokument direkt auf dieser Seite im Online-Reader ansehen und dann entscheiden, ob Sie es herunterladen möchten.
Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
Vorschau wird generiert...
Trích nội dung tài liệu
Pig Optimization and Execution Alan F. Gates @alanfgates © Hortonworks Inc. 2011 Page 1 Who Am I? Pig committer and PMC Member HCatalog committer and mentor Member of ASF and Incubator PMC Co-founder of Hortonworks Author of Programming Pig from O’Reilly Photo credit: Steven Guarnaccia, The Three Little Pigs Who Are You? 3 What Should We Optimize? Minimize scans – Hadoop is still often I/O bound Minimize total number of MR jobs Minimize shuffle size and number of shuffles Avoid spills to disk Reduce or remove skew For small jobs, minimize start-up time 4 Pig Deployment No server, all optimization and planning done on the launching machine Job executes on cluster Pig resides on user machine or gateway Hadoop Cluster User machine Pig Guts (i.e. Pig Architecture), p. 1 Logical Plan Pig Latin Load A = LOAD ‘myfile’ AS (x, y, z); B = GROUP A by x; C = FILTER B by group > 0; D = FOREACH C GENERATE group, COUNT(A); STORE D INTO ‘output’; Group AST Filter Foreach Semantic Checks Store 6 Pig Guts, p. 2 MapReduce Plan Logical Plan Load Load Group Filter Filter Foreach Map Filter Rearrange Group Rule based optimizations Reduce Store 7 Foreach Package Store Foreach Pig Guts, p. 3 MapReduce Plan Map Map Filter Filter Rearrange Rearrange Combine Foreach Physical optimizations Reduce Package Reduce Package Foreach Foreach 8 It would be really cool if… Map Map Reduce Reduce Map Map Reduce Reduce Map Reduce What’s the right join algorithm here? Even with statistics it would be hard to know. Need on the fly execution plan rewrites. 9 Memory Java + memory management = oil + water Java types inefficient memory users (~4x disk size) Very difficult to tell how much memory you are using Originally tried to monitor memory use via MXBeans: FAIL! Now estimate number of records we can hold in memory and spill when we exceed; allow user to tune guess 10 Reducing Spills to Disk Select Map size an
- Dokumentenname
- Pig II (11) (Tối ưu hóa và thực thi Pig II) (Tiếng Anh)
- Schule / Kurs
- Duke University · Big Data
- Autor (im Dokument)
- Alan F. Gates
- Inhalt
- Tài liệu này cung cấp cái nhìn sâu sắc về kiến trúc và các phương pháp tối ưu hóa hiệu suất cho Apache Pig, bao gồm giảm thiểu I/O, số lượng job MapReduce, và xử lý dữ liệu lệch. Nó cũng thảo luận về quản lý bộ nhớ, khởi động job nhanh hơn và các mô hình thực thi cải tiến.
- Inhaltsverzeichnis
- Who Am I?
- Who Are You?
- What Should We Optimize?
- Pig Deployment
- Pig Guts (i.e. Pig Architecture), p. 1
- Pig Guts, p. 2
- Pig Guts, p. 3
- It would be really cool if…
- Memory
- Reducing Spills to Disk
- Skew
- Reducing your Reducers
- (De)serialization
- Faster Job Startup
- Improved Execution Models
- Code Generation
- Learn More
- Questions?
- Seiten
- 19 Seiten
- Hochgeladen von
- Uni24h
Kommentare (0)
Noch keine Kommentare. Seien Sie der Erste!
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 10)
Advanced Big Data Analytics - Phân tích dữ liệu lớn nâng cao (Lecture 6)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 3)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 4)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 5)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

Kommentare (0)
Noch keine Kommentare. Seien Sie der Erste!