Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
- Seiten
- 54
- Định dạng
- PPT
- Dung lượng
- 2.5 MB
- Trường
- Duke University
- Aufrufe
- 0
- Kommentare
- 0
- Lượt tải
- 0
Vorschau wird generiert...
Bài giảng giới thiệu về Pig, một hệ thống xử lý dữ liệu bậc cao trên Hadoop, so sánh Pig Latin với MapReduce và trình bày các khái niệm cơ bản về toán tử, kiểu dữ liệu và schema.
- Dokumentenname
- Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
- Schule / Kurs
- Duke University · Big Data
- Inhalt
- Tài liệu giới thiệu Pig như một giải pháp thay thế MapReduce hiệu quả hơn cho xử lý dữ liệu lớn trên Hadoop. Nó tập trung vào ngôn ngữ Pig Latin, các khái niệm cốt lõi và lợi ích so với MapReduce.
- Inhaltsverzeichnis
- Is MapReduce not Good Enough?
- Pig to the Rescue
- How Pig is used in the Industry
- Motivation by Example
- In MapReduce
- In Pig Latin
- Pig runs over Hadoop
- Wait a minute
- MapReduce Vs. Pig cont.
- Pig Latin
- Statement
- Schema
- Schema cont.
- Data Types
- Data Types cont.
- Date Types cont.
- Operators
- Functions
- Seiten
- 54 Seiten
- Hochgeladen von
- Uni24h
Beschreibung
Trích nội dung tài liệu
Pig, a high level data processing system on Hadoop Is MapReduce not Good Enough? Restricted programming model Only two phases ◼ Job chain for long data flow Too many lines of code even for simple logic How many lines do you have for word count? ◼ Programmers are responsible for this Pig to the Rescue High level dataflow language (Pig Latin) Much simpler than Java ◼ Simplifies the data processing Puts the operations at the apropriate phases ◼ Chains multiple MR jobs How Pig is used in the Industry At Yahoo, 70% MapReduce jobs are written in Pig ◼ Used to Process web logs ◼ Build user behavior models ◼ Process images ◼ Data mining Also used by Twitter, LinkedIn, eBay, AOL, ... Motivation by Example Suppose we have user data in one file, website data in another file. ◼ We need to find the top 5 most visited pages by users aged 18-25 In MapReduce In Pig Latin Pig runs over Hadoop Wait a minute How to map the data to records By default, one line → one record User can customize the loading process How to identify attributes and map them to the schema Delimiter to separate different attributes By default, delimiter is tab. Customizable. MapReduce Vs. Pig cont. Join in MapReduce Various algorithms. None of them are easy to implement in MapReduce Multi-way join is more complicated Hard to integrate into SPJA workflow MapReduce Vs. Pig cont. Join in Pig Various algorithms are already available. Some of them are generic to support multi-way join No need to consider integration into SPJA workflow. Pig does that for you! A = LOAD 'input/join/A'; B = LOAD 'input/join/B'; C = JOIN A BY $0, B BY $1; DUMP C; Pig Latin Data flow language Users specify a sequence of operations to process data More control on the process, compared with declarative language Various data types are supported ◼ Schema is supported ◼ User-defined functions are supported
Häufig gestellte Fragen
Ist dieses Dokument kostenlos?
Ja. „Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)“ ist kostenlos — melden Sie sich einfach an und klicken Sie auf Herunterladen, um die Originaldatei zu erhalten.
Wie viele Seiten hat dieses Dokument?
Das Dokument hat 54 Seiten, für den Kurs Big Data. Sie können es vor dem Herunterladen online in der Vorschau ansehen.
Kann ich vor dem Herunterladen eine Vorschau ansehen?
Ja. Sie können sich dieses Dokument direkt auf dieser Seite im Online-Reader ansehen und dann entscheiden, ob Sie es herunterladen möchten.
Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
Vorschau wird generiert...
Trích nội dung tài liệu
Pig, a high level data processing system on Hadoop Is MapReduce not Good Enough? Restricted programming model Only two phases ◼ Job chain for long data flow Too many lines of code even for simple logic How many lines do you have for word count? ◼ Programmers are responsible for this Pig to the Rescue High level dataflow language (Pig Latin) Much simpler than Java ◼ Simplifies the data processing Puts the operations at the apropriate phases ◼ Chains multiple MR jobs How Pig is used in the Industry At Yahoo, 70% MapReduce jobs are written in Pig ◼ Used to Process web logs ◼ Build user behavior models ◼ Process images ◼ Data mining Also used by Twitter, LinkedIn, eBay, AOL, ... Motivation by Example Suppose we have user data in one file, website data in another file. ◼ We need to find the top 5 most visited pages by users aged 18-25 In MapReduce In Pig Latin Pig runs over Hadoop Wait a minute How to map the data to records By default, one line → one record User can customize the loading process How to identify attributes and map them to the schema Delimiter to separate different attributes By default, delimiter is tab. Customizable. MapReduce Vs. Pig cont. Join in MapReduce Various algorithms. None of them are easy to implement in MapReduce Multi-way join is more complicated Hard to integrate into SPJA workflow MapReduce Vs. Pig cont. Join in Pig Various algorithms are already available. Some of them are generic to support multi-way join No need to consider integration into SPJA workflow. Pig does that for you! A = LOAD 'input/join/A'; B = LOAD 'input/join/B'; C = JOIN A BY $0, B BY $1; DUMP C; Pig Latin Data flow language Users specify a sequence of operations to process data More control on the process, compared with declarative language Various data types are supported ◼ Schema is supported ◼ User-defined functions are supported
- Dokumentenname
- Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
- Schule / Kurs
- Duke University · Big Data
- Inhalt
- Tài liệu giới thiệu Pig như một giải pháp thay thế MapReduce hiệu quả hơn cho xử lý dữ liệu lớn trên Hadoop. Nó tập trung vào ngôn ngữ Pig Latin, các khái niệm cốt lõi và lợi ích so với MapReduce.
- Inhaltsverzeichnis
- Is MapReduce not Good Enough?
- Pig to the Rescue
- How Pig is used in the Industry
- Motivation by Example
- In MapReduce
- In Pig Latin
- Pig runs over Hadoop
- Wait a minute
- MapReduce Vs. Pig cont.
- Pig Latin
- Statement
- Schema
- Schema cont.
- Data Types
- Data Types cont.
- Date Types cont.
- Operators
- Functions
- Seiten
- 54 Seiten
- Hochgeladen von
- Uni24h
Kommentare (0)
Noch keine Kommentare. Seien Sie der Erste!
Stream (11) (Xử lý luồng dữ liệu) - Julian M. Kunkel
Krone (09) (Sự phát triển của dữ liệu) (Tiếng Anh)
Parallel mf (09) (Thuật toán phân tán phân tích ma trận dữ liệu lớn)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 5)
NoSQL db (06) (Cơ sở dữ liệu NoSQL)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

Kommentare (0)
Noch keine Kommentare. Seien Sie der Erste!