Giới thiệu về Data Flow Languages và Apache Pig - Julian M. Kunkel
Génération de l'aperçu...
Slide bài giảng giới thiệu về Data Flow Languages và Apache Pig, bao gồm mô hình dữ liệu, ngôn ngữ Pig Latin, truy cập dữ liệu và kiến trúc thực thi.
Description
Data Flow Languages & Apache Pig Lecture BigData Analytics Julian M. Kunkel julian.kunkel@googlemail.com University of Hamburg / German Climate Computing Center (DKRZ) 2017-01-13 Disclaimer: Big Data software is constantly updated, code samples may be outdated. Overview Pig Latin Accessing Data Architecture Summary Outline 1 Overview 2 Pig Latin 3 Accessing Data 4 Architecture 5 Summary Julian M. Kunkel Lecture BigData Analytics, 2016 2 / 32 Overview Pig Latin Accessing Data Architecture Summary General Data Model for Dataflow Languages Data Tuple t = (x1 , ..., xn ) where xi may be of a given type Input/Output = list of tuples (like a table) Typical Operators for Data-Flow Processing Operations process individual tuples Map/Foreach: process or transform data of individual tuples or group transform a tuple: student.Map((matrikel, name) ⇒ (matrikel + 4, name)) count members for each group: groupedStudents.Map((year) ⇒ count()) Filter tuples by comparing a key to a value Operations that require the complete input data Group tuples by a key Sort data according to a key Join multiple relations together Split tuples of a relation into multiple relations (based on a condition) Julian M. Kunkel Lecture BigData Analytics, 2016 3 / 32 Overview Pig Latin Accessing Data Architecture Summary Data Flow Programming Paradigm [68] Focus: data movement and transformation Compare to imperative programming: sequence of commands Models program as directed graph of data flowing between operations Input/output is illustrated as a node Node is an operation, edges are dependencies Operation is run once all inputs become valid An operation might work on a single data element or on the complete data Parallelism is inherently supported by data flow languages States (in the program) Dataflow works best with stateless programs Stateful dataflow graphs support mutable state Data related states, e.g., reductions, may be encoded as data Programming Func
Résumé IA
- Nom du document
- Giới thiệu về Data Flow Languages và Apache Pig - Julian M. Kunkel
- École / Cours
- University of Hamburg · Big Data
- Contenu
- Tài liệu này giới thiệu về các khái niệm cơ bản của ngôn ngữ luồng dữ liệu và Apache Pig, một nền tảng xử lý dữ liệu lớn. Nó mô tả cách dữ liệu được biểu diễn, các toán tử xử lý, mô hình lập trình, và cách trực quan hóa bằng sơ đồ luồng, cũng như kiến trúc và ngôn ngữ Pig Latin của Apache Pig.
- Table des matières
- Overview
- Pig Latin
- Accessing Data
- Architecture
- Summary
- Pages
- 33 pages
- Téléversé par
- Uni24h
Foire aux questions
Ce document est-il gratuit ?
Oui. « Giới thiệu về Data Flow Languages và Apache Pig - Julian M. Kunkel » est gratuit — il suffit de vous connecter et de cliquer sur Télécharger pour obtenir le fichier original.
Combien de pages compte ce document ?
Le document contient 33 pages, pour le cours Big Data. Vous pouvez le prévisualiser en ligne avant de le télécharger.
Puis-je prévisualiser avant de télécharger ?
Oui. Vous pouvez prévisualiser ce document directement sur cette page avec le lecteur en ligne, puis décider de le télécharger ou non.
Giới thiệu về Data Flow Languages và Apache Pig - Julian M. Kunkel
Génération de l'aperçu...
Data Flow Languages & Apache Pig Lecture BigData Analytics Julian M. Kunkel julian.kunkel@googlemail.com University of Hamburg / German Climate Computing Center (DKRZ) 2017-01-13 Disclaimer: Big Data software is constantly updated, code samples may be outdated. Overview Pig Latin Accessing Data Architecture Summary Outline 1 Overview 2 Pig Latin 3 Accessing Data 4 Architecture 5 Summary Julian M. Kunkel Lecture BigData Analytics, 2016 2 / 32 Overview Pig Latin Accessing Data Architecture Summary General Data Model for Dataflow Languages Data Tuple t = (x1 , ..., xn ) where xi may be of a given type Input/Output = list of tuples (like a table) Typical Operators for Data-Flow Processing Operations process individual tuples Map/Foreach: process or transform data of individual tuples or group transform a tuple: student.Map((matrikel, name) ⇒ (matrikel + 4, name)) count members for each group: groupedStudents.Map((year) ⇒ count()) Filter tuples by comparing a key to a value Operations that require the complete input data Group tuples by a key Sort data according to a key Join multiple relations together Split tuples of a relation into multiple relations (based on a condition) Julian M. Kunkel Lecture BigData Analytics, 2016 3 / 32 Overview Pig Latin Accessing Data Architecture Summary Data Flow Programming Paradigm [68] Focus: data movement and transformation Compare to imperative programming: sequence of commands Models program as directed graph of data flowing between operations Input/output is illustrated as a node Node is an operation, edges are dependencies Operation is run once all inputs become valid An operation might work on a single data element or on the complete data Parallelism is inherently supported by data flow languages States (in the program) Dataflow works best with stateless programs Stateful dataflow graphs support mutable state Data related states, e.g., reductions, may be encoded as data Programming Func
Lire le document entier
- Nom du document
- Giới thiệu về Data Flow Languages và Apache Pig - Julian M. Kunkel
- École / Cours
- University of Hamburg · Big Data
- Contenu
- Tài liệu này giới thiệu về các khái niệm cơ bản của ngôn ngữ luồng dữ liệu và Apache Pig, một nền tảng xử lý dữ liệu lớn. Nó mô tả cách dữ liệu được biểu diễn, các toán tử xử lý, mô hình lập trình, và cách trực quan hóa bằng sơ đồ luồng, cũng như kiến trúc và ngôn ngữ Pig Latin của Apache Pig.
- Table des matières
- Overview
- Pig Latin
- Accessing Data
- Architecture
- Summary
- Pages
- 33 pages
- Téléversé par
- Uni24h
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !