Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
Génération de l'aperçu...
Bài giảng giới thiệu về Pig, một hệ thống xử lý dữ liệu bậc cao trên Hadoop, so sánh Pig Latin với MapReduce và trình bày các khái niệm cơ bản về toán tử, kiểu dữ liệu và schema.
Description
Pig, a high level data processing system on Hadoop Is MapReduce not Good Enough? Restricted programming model Only two phases ◼ Job chain for long data flow Too many lines of code even for simple logic How many lines do you have for word count? ◼ Programmers are responsible for this Pig to the Rescue High level dataflow language (Pig Latin) Much simpler than Java ◼ Simplifies the data processing Puts the operations at the apropriate phases ◼ Chains multiple MR jobs How Pig is used in the Industry At Yahoo, 70% MapReduce jobs are written in Pig ◼ Used to Process web logs ◼ Build user behavior models ◼ Process images ◼ Data mining Also used by Twitter, LinkedIn, eBay, AOL, ... Motivation by Example Suppose we have user data in one file, website data in another file. ◼ We need to find the top 5 most visited pages by users aged 18-25 In MapReduce In Pig Latin Pig runs over Hadoop Wait a minute How to map the data to records By default, one line → one record User can customize the loading process How to identify attributes and map them to the schema Delimiter to separate different attributes By default, delimiter is tab. Customizable. MapReduce Vs. Pig cont. Join in MapReduce Various algorithms. None of them are easy to implement in MapReduce Multi-way join is more complicated Hard to integrate into SPJA workflow MapReduce Vs. Pig cont. Join in Pig Various algorithms are already available. Some of them are generic to support multi-way join No need to consider integration into SPJA workflow. Pig does that for you! A = LOAD 'input/join/A'; B = LOAD 'input/join/B'; C = JOIN A BY $0, B BY $1; DUMP C; Pig Latin Data flow language Users specify a sequence of operations to process data More control on the process, compared with declarative language Various data types are supported ◼ Schema is supported ◼ User-defined functions are supported
Résumé IA
- Nom du document
- Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
- École / Cours
- Duke University · Big Data
- Contenu
- Tài liệu giới thiệu Pig như một giải pháp thay thế MapReduce hiệu quả hơn cho xử lý dữ liệu lớn trên Hadoop. Nó tập trung vào ngôn ngữ Pig Latin, các khái niệm cốt lõi và lợi ích so với MapReduce.
- Table des matières
- Is MapReduce not Good Enough?
- Pig to the Rescue
- How Pig is used in the Industry
- Motivation by Example
- In MapReduce
- In Pig Latin
- Pig runs over Hadoop
- Wait a minute
- MapReduce Vs. Pig cont.
- Pig Latin
- Statement
- Schema
- Schema cont.
- Data Types
- Data Types cont.
- Date Types cont.
- Operators
- Functions
- Pages
- 54 pages
- Téléversé par
- Uni24h
Foire aux questions
Ce document est-il gratuit ?
Oui. « Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop) » est gratuit — il suffit de vous connecter et de cliquer sur Télécharger pour obtenir le fichier original.
Combien de pages compte ce document ?
Le document contient 54 pages, pour le cours Big Data. Vous pouvez le prévisualiser en ligne avant de le télécharger.
Puis-je prévisualiser avant de télécharger ?
Oui. Vous pouvez prévisualiser ce document directement sur cette page avec le lecteur en ligne, puis décider de le télécharger ou non.
Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
Génération de l'aperçu...
Pig, a high level data processing system on Hadoop Is MapReduce not Good Enough? Restricted programming model Only two phases ◼ Job chain for long data flow Too many lines of code even for simple logic How many lines do you have for word count? ◼ Programmers are responsible for this Pig to the Rescue High level dataflow language (Pig Latin) Much simpler than Java ◼ Simplifies the data processing Puts the operations at the apropriate phases ◼ Chains multiple MR jobs How Pig is used in the Industry At Yahoo, 70% MapReduce jobs are written in Pig ◼ Used to Process web logs ◼ Build user behavior models ◼ Process images ◼ Data mining Also used by Twitter, LinkedIn, eBay, AOL, ... Motivation by Example Suppose we have user data in one file, website data in another file. ◼ We need to find the top 5 most visited pages by users aged 18-25 In MapReduce In Pig Latin Pig runs over Hadoop Wait a minute How to map the data to records By default, one line → one record User can customize the loading process How to identify attributes and map them to the schema Delimiter to separate different attributes By default, delimiter is tab. Customizable. MapReduce Vs. Pig cont. Join in MapReduce Various algorithms. None of them are easy to implement in MapReduce Multi-way join is more complicated Hard to integrate into SPJA workflow MapReduce Vs. Pig cont. Join in Pig Various algorithms are already available. Some of them are generic to support multi-way join No need to consider integration into SPJA workflow. Pig does that for you! A = LOAD 'input/join/A'; B = LOAD 'input/join/B'; C = JOIN A BY $0, B BY $1; DUMP C; Pig Latin Data flow language Users specify a sequence of operations to process data More control on the process, compared with declarative language Various data types are supported ◼ Schema is supported ◼ User-defined functions are supported
Lire le document entier
- Nom du document
- Pig I (Hệ thống xử lý dữ liệu bậc cao trên Hadoop)
- École / Cours
- Duke University · Big Data
- Contenu
- Tài liệu giới thiệu Pig như một giải pháp thay thế MapReduce hiệu quả hơn cho xử lý dữ liệu lớn trên Hadoop. Nó tập trung vào ngôn ngữ Pig Latin, các khái niệm cốt lõi và lợi ích so với MapReduce.
- Table des matières
- Is MapReduce not Good Enough?
- Pig to the Rescue
- How Pig is used in the Industry
- Motivation by Example
- In MapReduce
- In Pig Latin
- Pig runs over Hadoop
- Wait a minute
- MapReduce Vs. Pig cont.
- Pig Latin
- Statement
- Schema
- Schema cont.
- Data Types
- Data Types cont.
- Date Types cont.
- Operators
- Functions
- Pages
- 54 pages
- Téléversé par
- Uni24h
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !