How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
正在生成预览...
Tài liệu slide bài giảng về cách MapReduce hoạt động trong Hadoop, bao gồm vòng đời tác vụ, kiến trúc HDFS, xử lý lỗi và tối ưu tham số cấu hình.
描述
Data-intensive Computing Systems How MapReduce Works (in Hadoop) Shivnath Babu Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 Components in a Hadoop MR Workflow Next few slides are from: http://www.slideshare.net/hadoop/practical-problem-solving-with-apache-hadoop-pig Job Submission Initialization Scheduling Execution Map Task Sort Buffer Reduce Tasks Quick Overview of Other Topics (Will Revisit Them Later in the Course) Dealing with failures Hadoop Distributed FileSystem (HDFS) Optimizing a MapReduce job Dealing with Failures and Slow Tasks What to do when a task fails? Try again (retries possible because of idempotence) Try again somewhere else Report failure What about slow tasks: stragglers Run another version of the same task in parallel. Take results from the one that finishes first What are the pros and cons of this approach? Fault tolerance is of high priority in the MapReduce framework HDFS Architecture Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 How are the number of splits, number of map and reduce tasks, memory allocation to tasks, etc., determined? Job Configuration Parameters 190+ parameters in Hadoop Set manually or defaults are used Hadoop Job Configuration Parameters Image source: http://www.jaso.co.kr/265 Tuning Hadoop Job Conf. Parameters Do their settings impact performance? What are ways to set these parameters? Defaults -- are they good enough? Best practices -- the best setting can depend on data, job, and cluster properties Automatic setting Experimental Setting Hadoop cluster on 1 master + 16 workers Each node: 2GHz AMD processor, 1.8GB RAM, 30GB local disk
AI 摘要
- 文档名称
- How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
- 学校 / 课程
- Duke University · Big Data
- 内容
- Tài liệu giải thích chi tiết cách MapReduce hoạt động trong Hadoop, từ vòng đời công việc đến các thành phần và xử lý lỗi. Nó cũng đề cập đến HDFS và các tham số cấu hình, cùng với các thử nghiệm thực nghiệm.
- 目录
- Lifecycle of a MapReduce Job
- Components in a Hadoop MR Workflow
- Job Submission
- Initialization
- Scheduling
- Execution
- Map Task
- Sort Buffer
- Reduce Tasks
- Quick Overview of Other Topics (Will Revisit Them Later in the Course)
- Dealing with Failures and Slow Tasks
- HDFS Architecture
- Lifecycle of a MapReduce Job
- Job Configuration Parameters
- Hadoop Job Configuration Parameters
- Tuning Hadoop Job Conf. Parameters
- Experimental Setting
- Parameters Varied in Experiments
- Hadoop 50GB TeraSort
- Hadoop 75GB TeraSort
- Automatic Optimization? (Not yet in Hadoop)
- 页数
- 26 页
- 上传者
- Uni24h
常见问题
此文档免费吗?
是的。“How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)”是免费的 — 只需登录并点击“下载”即可获取原始文件。
这份文档有多少页?
该文档共有 26 页,适用于课程 Big Data。您可以在下载前进行在线预览。
我可以在下载前预览吗?
是的。您可以通过在线阅读器直接在本页面预览此文档,然后再决定是否下载。
How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
正在生成预览...
Data-intensive Computing Systems How MapReduce Works (in Hadoop) Shivnath Babu Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 Components in a Hadoop MR Workflow Next few slides are from: http://www.slideshare.net/hadoop/practical-problem-solving-with-apache-hadoop-pig Job Submission Initialization Scheduling Execution Map Task Sort Buffer Reduce Tasks Quick Overview of Other Topics (Will Revisit Them Later in the Course) Dealing with failures Hadoop Distributed FileSystem (HDFS) Optimizing a MapReduce job Dealing with Failures and Slow Tasks What to do when a task fails? Try again (retries possible because of idempotence) Try again somewhere else Report failure What about slow tasks: stragglers Run another version of the same task in parallel. Take results from the one that finishes first What are the pros and cons of this approach? Fault tolerance is of high priority in the MapReduce framework HDFS Architecture Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 How are the number of splits, number of map and reduce tasks, memory allocation to tasks, etc., determined? Job Configuration Parameters 190+ parameters in Hadoop Set manually or defaults are used Hadoop Job Configuration Parameters Image source: http://www.jaso.co.kr/265 Tuning Hadoop Job Conf. Parameters Do their settings impact performance? What are ways to set these parameters? Defaults -- are they good enough? Best practices -- the best setting can depend on data, job, and cluster properties Automatic setting Experimental Setting Hadoop cluster on 1 master + 16 workers Each node: 2GHz AMD processor, 1.8GB RAM, 30GB local disk
阅读全文
- 文档名称
- How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
- 学校 / 课程
- Duke University · Big Data
- 内容
- Tài liệu giải thích chi tiết cách MapReduce hoạt động trong Hadoop, từ vòng đời công việc đến các thành phần và xử lý lỗi. Nó cũng đề cập đến HDFS và các tham số cấu hình, cùng với các thử nghiệm thực nghiệm.
- 目录
- Lifecycle of a MapReduce Job
- Components in a Hadoop MR Workflow
- Job Submission
- Initialization
- Scheduling
- Execution
- Map Task
- Sort Buffer
- Reduce Tasks
- Quick Overview of Other Topics (Will Revisit Them Later in the Course)
- Dealing with Failures and Slow Tasks
- HDFS Architecture
- Lifecycle of a MapReduce Job
- Job Configuration Parameters
- Hadoop Job Configuration Parameters
- Tuning Hadoop Job Conf. Parameters
- Experimental Setting
- Parameters Varied in Experiments
- Hadoop 50GB TeraSort
- Hadoop 75GB TeraSort
- Automatic Optimization? (Not yet in Hadoop)
- 页数
- 26 页
- 上传者
- Uni24h
评论 (0)
暂无评论。快来抢沙发吧!
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
评论 (0)
暂无评论。快来抢沙发吧!