Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
Generating preview...
Bài giảng giới thiệu về Pig, một hệ thống xử lý dữ liệu bậc cao trên Hadoop, so sánh Pig Latin với MapReduce và trình bày các khái niệm cơ bản về toán tử, kiểu dữ liệu và schema.
Description
Pig, a high level data processing system on Hadoop Is MapReduce not Good Enough? Restricted programming model Only two phases ◼ Job chain for long data flow Too many lines of code even for simple logic How many lines do you have for word count? ◼ Programmers are responsible for this Pig to the Rescue High level dataflow language (Pig Latin) Much simpler than Java ◼ Simplifies the data processing Puts the operations at the apropriate phases ◼ Chains multiple MR jobs How Pig is used in the Industry At Yahoo, 70% MapReduce jobs are written in Pig ◼ Used to Process web logs ◼ Build user behavior models ◼ Process images ◼ Data mining Also used by Twitter, LinkedIn, eBay, AOL, ... Motivation by Example Suppose we have user data in one file, website data in another file. ◼ We need to find the top 5 most visited pages by users aged 18-25 In MapReduce In Pig Latin Pig runs over Hadoop Wait a minute How to map the data to records By default, one line → one record User can customize the loading process How to identify attributes and map them to the schema Delimiter to separate different attributes By default, delimiter is tab. Customizable. MapReduce Vs. Pig cont. Join in MapReduce Various algorithms. None of them are easy to implement in MapReduce Multi-way join is more complicated Hard to integrate into SPJA workflow MapReduce Vs. Pig cont. Join in Pig Various algorithms are already available. Some of them are generic to support multi-way join No need to consider integration into SPJA workflow. Pig does that for you! A = LOAD 'input/join/A'; B = LOAD 'input/join/B'; C = JOIN A BY $0, B BY $1; DUMP C; Pig Latin Data flow language Users specify a sequence of operations to process data More control on the process, compared with declarative language Various data types are supported ◼ Schema is supported ◼ User-defined functions are supported
AI summary
- Document name
- Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
- School / Course
- Duke University · Big Data
- Content
- Tài liệu giới thiệu Pig như một giải pháp thay thế MapReduce hiệu quả hơn cho xử lý dữ liệu lớn trên Hadoop. Nó tập trung vào ngôn ngữ Pig Latin, các khái niệm cốt lõi và lợi ích so với MapReduce.
- Table of contents
- Is MapReduce not Good Enough?
- Pig to the Rescue
- How Pig is used in the Industry
- Motivation by Example
- In MapReduce
- In Pig Latin
- Pig runs over Hadoop
- Wait a minute
- MapReduce Vs. Pig cont.
- Pig Latin
- Statement
- Schema
- Schema cont.
- Data Types
- Data Types cont.
- Date Types cont.
- Operators
- Functions
- Pages
- 54 pages
- Uploaded by
- Uni24h
Frequently asked questions
Is this document free?
Yes. “Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)” is free — just sign in and click Download to get the original file.
How many pages is this document?
The document has 54 pages, for the course Big Data. You can preview it online before downloading.
Can I preview before downloading?
Yes. You can preview this document right on this page with the online reader, then decide whether to download.
Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
Generating preview...
Pig, a high level data processing system on Hadoop Is MapReduce not Good Enough? Restricted programming model Only two phases ◼ Job chain for long data flow Too many lines of code even for simple logic How many lines do you have for word count? ◼ Programmers are responsible for this Pig to the Rescue High level dataflow language (Pig Latin) Much simpler than Java ◼ Simplifies the data processing Puts the operations at the apropriate phases ◼ Chains multiple MR jobs How Pig is used in the Industry At Yahoo, 70% MapReduce jobs are written in Pig ◼ Used to Process web logs ◼ Build user behavior models ◼ Process images ◼ Data mining Also used by Twitter, LinkedIn, eBay, AOL, ... Motivation by Example Suppose we have user data in one file, website data in another file. ◼ We need to find the top 5 most visited pages by users aged 18-25 In MapReduce In Pig Latin Pig runs over Hadoop Wait a minute How to map the data to records By default, one line → one record User can customize the loading process How to identify attributes and map them to the schema Delimiter to separate different attributes By default, delimiter is tab. Customizable. MapReduce Vs. Pig cont. Join in MapReduce Various algorithms. None of them are easy to implement in MapReduce Multi-way join is more complicated Hard to integrate into SPJA workflow MapReduce Vs. Pig cont. Join in Pig Various algorithms are already available. Some of them are generic to support multi-way join No need to consider integration into SPJA workflow. Pig does that for you! A = LOAD 'input/join/A'; B = LOAD 'input/join/B'; C = JOIN A BY $0, B BY $1; DUMP C; Pig Latin Data flow language Users specify a sequence of operations to process data More control on the process, compared with declarative language Various data types are supported ◼ Schema is supported ◼ User-defined functions are supported
Read full document
- Document name
- Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
- School / Course
- Duke University · Big Data
- Content
- Tài liệu giới thiệu Pig như một giải pháp thay thế MapReduce hiệu quả hơn cho xử lý dữ liệu lớn trên Hadoop. Nó tập trung vào ngôn ngữ Pig Latin, các khái niệm cốt lõi và lợi ích so với MapReduce.
- Table of contents
- Is MapReduce not Good Enough?
- Pig to the Rescue
- How Pig is used in the Industry
- Motivation by Example
- In MapReduce
- In Pig Latin
- Pig runs over Hadoop
- Wait a minute
- MapReduce Vs. Pig cont.
- Pig Latin
- Statement
- Schema
- Schema cont.
- Data Types
- Data Types cont.
- Date Types cont.
- Operators
- Functions
- Pages
- 54 pages
- Uploaded by
- Uni24h
Comments (0)
No comments yet. Be the first!
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Comments (0)
No comments yet. Be the first!