How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
- 페이지 수
- 26
- 형식
- PPT
- 크기
- 5.9 MB
- Trường
- Duke University
- 조회수
- 0
- 댓글
- 0
- Lượt tải
- 0
미리보기 생성 중...
Tài liệu slide bài giảng về cách MapReduce hoạt động trong Hadoop, bao gồm vòng đời tác vụ, kiến trúc HDFS, xử lý lỗi và tối ưu tham số cấu hình.
- 문서명
- How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
- 학교 / 강의
- Duke University · Big Data
- 내용
- Tài liệu giải thích chi tiết cách MapReduce hoạt động trong Hadoop, từ vòng đời công việc đến các thành phần và xử lý lỗi. Nó cũng đề cập đến HDFS và các tham số cấu hình, cùng với các thử nghiệm thực nghiệm.
- 목차
- Lifecycle of a MapReduce Job
- Components in a Hadoop MR Workflow
- Job Submission
- Initialization
- Scheduling
- Execution
- Map Task
- Sort Buffer
- Reduce Tasks
- Quick Overview of Other Topics (Will Revisit Them Later in the Course)
- Dealing with Failures and Slow Tasks
- HDFS Architecture
- Lifecycle of a MapReduce Job
- Job Configuration Parameters
- Hadoop Job Configuration Parameters
- Tuning Hadoop Job Conf. Parameters
- Experimental Setting
- Parameters Varied in Experiments
- Hadoop 50GB TeraSort
- Hadoop 75GB TeraSort
- Automatic Optimization? (Not yet in Hadoop)
- 페이지 수
- 26 페이지
- 업로더
- Uni24h
설명
Trích nội dung tài liệu
Data-intensive Computing Systems How MapReduce Works (in Hadoop) Shivnath Babu Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 Components in a Hadoop MR Workflow Next few slides are from: http://www.slideshare.net/hadoop/practical-problem-solving-with-apache-hadoop-pig Job Submission Initialization Scheduling Execution Map Task Sort Buffer Reduce Tasks Quick Overview of Other Topics (Will Revisit Them Later in the Course) Dealing with failures Hadoop Distributed FileSystem (HDFS) Optimizing a MapReduce job Dealing with Failures and Slow Tasks What to do when a task fails? Try again (retries possible because of idempotence) Try again somewhere else Report failure What about slow tasks: stragglers Run another version of the same task in parallel. Take results from the one that finishes first What are the pros and cons of this approach? Fault tolerance is of high priority in the MapReduce framework HDFS Architecture Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 How are the number of splits, number of map and reduce tasks, memory allocation to tasks, etc., determined? Job Configuration Parameters 190+ parameters in Hadoop Set manually or defaults are used Hadoop Job Configuration Parameters Image source: http://www.jaso.co.kr/265 Tuning Hadoop Job Conf. Parameters Do their settings impact performance? What are ways to set these parameters? Defaults -- are they good enough? Best practices -- the best setting can depend on data, job, and cluster properties Automatic setting Experimental Setting Hadoop cluster on 1 master + 16 workers Each node: 2GHz AMD processor, 1.8GB RAM, 30GB local disk
자주 묻는 질문
이 문서는 무료인가요?
네. “How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)” 문서는 무료입니다. 로그인 후 '다운로드'를 클릭하여 원본 파일을 받으세요.
이 문서는 몇 페이지로 되어 있나요?
이 문서는 26페이지입니다, Big Data 과정용. 다운로드하기 전에 온라인으로 미리 볼 수 있습니다.
다운로드하기 전에 미리 볼 수 있나요?
네. 이 페이지의 온라인 리더를 통해 문서를 미리 본 후 다운로드 여부를 결정할 수 있습니다.
How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
미리보기 생성 중...
Trích nội dung tài liệu
Data-intensive Computing Systems How MapReduce Works (in Hadoop) Shivnath Babu Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Map function Reduce function Run this program as a MapReduce job Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 Components in a Hadoop MR Workflow Next few slides are from: http://www.slideshare.net/hadoop/practical-problem-solving-with-apache-hadoop-pig Job Submission Initialization Scheduling Execution Map Task Sort Buffer Reduce Tasks Quick Overview of Other Topics (Will Revisit Them Later in the Course) Dealing with failures Hadoop Distributed FileSystem (HDFS) Optimizing a MapReduce job Dealing with Failures and Slow Tasks What to do when a task fails? Try again (retries possible because of idempotence) Try again somewhere else Report failure What about slow tasks: stragglers Run another version of the same task in parallel. Take results from the one that finishes first What are the pros and cons of this approach? Fault tolerance is of high priority in the MapReduce framework HDFS Architecture Lifecycle of a MapReduce Job Time Input Splits Map Wave 1 Map Wave 2 Reduce Wave 1 Reduce Wave 2 How are the number of splits, number of map and reduce tasks, memory allocation to tasks, etc., determined? Job Configuration Parameters 190+ parameters in Hadoop Set manually or defaults are used Hadoop Job Configuration Parameters Image source: http://www.jaso.co.kr/265 Tuning Hadoop Job Conf. Parameters Do their settings impact performance? What are ways to set these parameters? Defaults -- are they good enough? Best practices -- the best setting can depend on data, job, and cluster properties Automatic setting Experimental Setting Hadoop cluster on 1 master + 16 workers Each node: 2GHz AMD processor, 1.8GB RAM, 30GB local disk
- 문서명
- How hadoop works (Cách MapReduce hoạt động trong Hadoop) (Tiếng Anh)
- 학교 / 강의
- Duke University · Big Data
- 내용
- Tài liệu giải thích chi tiết cách MapReduce hoạt động trong Hadoop, từ vòng đời công việc đến các thành phần và xử lý lỗi. Nó cũng đề cập đến HDFS và các tham số cấu hình, cùng với các thử nghiệm thực nghiệm.
- 목차
- Lifecycle of a MapReduce Job
- Components in a Hadoop MR Workflow
- Job Submission
- Initialization
- Scheduling
- Execution
- Map Task
- Sort Buffer
- Reduce Tasks
- Quick Overview of Other Topics (Will Revisit Them Later in the Course)
- Dealing with Failures and Slow Tasks
- HDFS Architecture
- Lifecycle of a MapReduce Job
- Job Configuration Parameters
- Hadoop Job Configuration Parameters
- Tuning Hadoop Job Conf. Parameters
- Experimental Setting
- Parameters Varied in Experiments
- Hadoop 50GB TeraSort
- Hadoop 75GB TeraSort
- Automatic Optimization? (Not yet in Hadoop)
- 페이지 수
- 26 페이지
- 업로더
- Uni24h
댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!
Stream (11) (Xử lý luồng dữ liệu) - Julian M. Kunkel
Krone (09) (Sự phát triển của dữ liệu) (Tiếng Anh)
Parallel mf (09) (Thuật toán phân tán phân tích ma trận dữ liệu lớn)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 5)
NoSQL db (06) (Cơ sở dữ liệu NoSQL)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!