How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
Génération de l'aperçu...
Tài liệu slide bài giảng về cách MapReduce hoạt động trong Hadoop, bao gồm vòng đời tác vụ, kiến trúc HDFS, xử lý lỗi và tối ưu tham số cấu hình.
Description
Data-intensive Computing Systems How MapReduce Works (in Hadoop) Shivnath Babu Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 Components in a Hadoop MR Workflow Next few slides are from: http://www.slideshare.net/hadoop/practical-problem-solving-with-apache-hadoop-pig Job Submission Initialization Scheduling Execution Map Task Sort Buffer Reduce Tasks Quick Overview of Other Topics (Will Revisit Them Later in the Course) Dealing with failures Hadoop Distributed FileSystem (HDFS) Optimizing a MapReduce job Dealing with Failures and Slow Tasks What to do when a task fails? Try again (retries possible because of idempotence) Try again somewhere else Report failure What about slow tasks: stragglers Run another version of the same task in parallel. Take results from the one that finishes first What are the pros and cons of this approach? Fault tolerance is of high priority in the MapReduce framework HDFS Architecture Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 How are the number of splits, number of map and reduce tasks, memory allocation to tasks, etc., determined? Job Configuration Parameters 190+ parameters in Hadoop Set manually or defaults are used Hadoop Job Configuration Parameters Image source: http://www.jaso.co.kr/265 Tuning Hadoop Job Conf. Parameters Do their settings impact performance? What are ways to set these parameters? Defaults -- are they good enough? Best practices -- the best setting can depend on data, job, and cluster properties Automatic setting Experimental Setting Hadoop cluster on 1 master + 16 workers Each node: 2GHz AMD processor, 1.8GB RAM, 30GB local disk
Résumé IA
- Nom du document
- How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
- École / Cours
- Duke University · Big Data
- Contenu
- Tài liệu giải thích chi tiết cách MapReduce hoạt động trong Hadoop, từ vòng đời công việc đến các thành phần và xử lý lỗi. Nó cũng đề cập đến HDFS và các tham số cấu hình, cùng với các thử nghiệm thực nghiệm.
- Table des matières
- Lifecycle of a MapReduce Job
- Components in a Hadoop MR Workflow
- Job Submission
- Initialization
- Scheduling
- Execution
- Map Task
- Sort Buffer
- Reduce Tasks
- Quick Overview of Other Topics (Will Revisit Them Later in the Course)
- Dealing with Failures and Slow Tasks
- HDFS Architecture
- Lifecycle of a MapReduce Job
- Job Configuration Parameters
- Hadoop Job Configuration Parameters
- Tuning Hadoop Job Conf. Parameters
- Experimental Setting
- Parameters Varied in Experiments
- Hadoop 50GB TeraSort
- Hadoop 75GB TeraSort
- Automatic Optimization? (Not yet in Hadoop)
- Pages
- 26 pages
- Téléversé par
- Uni24h
Foire aux questions
Ce document est-il gratuit ?
Oui. « How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh) » est gratuit — il suffit de vous connecter et de cliquer sur Télécharger pour obtenir le fichier original.
Combien de pages compte ce document ?
Le document contient 26 pages, pour le cours Big Data. Vous pouvez le prévisualiser en ligne avant de le télécharger.
Puis-je prévisualiser avant de télécharger ?
Oui. Vous pouvez prévisualiser ce document directement sur cette page avec le lecteur en ligne, puis décider de le télécharger ou non.
How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
Génération de l'aperçu...
Data-intensive Computing Systems How MapReduce Works (in Hadoop) Shivnath Babu Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 Components in a Hadoop MR Workflow Next few slides are from: http://www.slideshare.net/hadoop/practical-problem-solving-with-apache-hadoop-pig Job Submission Initialization Scheduling Execution Map Task Sort Buffer Reduce Tasks Quick Overview of Other Topics (Will Revisit Them Later in the Course) Dealing with failures Hadoop Distributed FileSystem (HDFS) Optimizing a MapReduce job Dealing with Failures and Slow Tasks What to do when a task fails? Try again (retries possible because of idempotence) Try again somewhere else Report failure What about slow tasks: stragglers Run another version of the same task in parallel. Take results from the one that finishes first What are the pros and cons of this approach? Fault tolerance is of high priority in the MapReduce framework HDFS Architecture Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 How are the number of splits, number of map and reduce tasks, memory allocation to tasks, etc., determined? Job Configuration Parameters 190+ parameters in Hadoop Set manually or defaults are used Hadoop Job Configuration Parameters Image source: http://www.jaso.co.kr/265 Tuning Hadoop Job Conf. Parameters Do their settings impact performance? What are ways to set these parameters? Defaults -- are they good enough? Best practices -- the best setting can depend on data, job, and cluster properties Automatic setting Experimental Setting Hadoop cluster on 1 master + 16 workers Each node: 2GHz AMD processor, 1.8GB RAM, 30GB local disk
Lire le document entier
- Nom du document
- How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
- École / Cours
- Duke University · Big Data
- Contenu
- Tài liệu giải thích chi tiết cách MapReduce hoạt động trong Hadoop, từ vòng đời công việc đến các thành phần và xử lý lỗi. Nó cũng đề cập đến HDFS và các tham số cấu hình, cùng với các thử nghiệm thực nghiệm.
- Table des matières
- Lifecycle of a MapReduce Job
- Components in a Hadoop MR Workflow
- Job Submission
- Initialization
- Scheduling
- Execution
- Map Task
- Sort Buffer
- Reduce Tasks
- Quick Overview of Other Topics (Will Revisit Them Later in the Course)
- Dealing with Failures and Slow Tasks
- HDFS Architecture
- Lifecycle of a MapReduce Job
- Job Configuration Parameters
- Hadoop Job Configuration Parameters
- Tuning Hadoop Job Conf. Parameters
- Experimental Setting
- Parameters Varied in Experiments
- Hadoop 50GB TeraSort
- Hadoop 75GB TeraSort
- Automatic Optimization? (Not yet in Hadoop)
- Pages
- 26 pages
- Téléversé par
- Uni24h
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !