Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)
- 페이지 수
- 79
- 형식
- 크기
- 10.2 MB
- 연도
- 2014
- Trường
- Columbia University
- 조회수
- 0
- 댓글
- 0
- Lượt tải
- 0
미리보기 생성 중...
Bài giảng số 2 về Nền tảng Phân tích Dữ liệu Lớn, tập trung vào Apache Hadoop và các dự án liên quan, cùng các trường hợp sử dụng và ví dụ về MapReduce.
- 문서명
- Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)
- 학교 / 강의
- Columbia University · Big Data
- 작성자 (문서 내)
- Ching-Yung Lin, Ph.D.
- 내용
- Giới thiệu về Apache Hadoop và các dự án liên quan
- 목차
- 이 문서는 명확한 목차가 없습니다.
- 페이지 수
- 79 페이지
- 업로더
- Uni24h
설명
Trích nội dung tài liệu
E6893 Big Data Analytics Lecture 2: Big Data Analytics Platforms Ching-Yung Lin, Ph.D. Adjunct Professor, Depts. of Electrical Engineering and Computer Science September 15th, 2023 1 E6893 Big Data Analytics – Lecture 2: Big Data Platform © 2023 CY Lin, Columbia University Remind -- Apache Hadoop The Apache™ Hadoop® project develops open-source software for reliable, scalable, distributed computing. The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures. The project includes these modules: Hadoop Common: The common utilities that support the other Hadoop modules. Hadoop Distributed File System (HDFS™): A distributed file system that provides highthroughput access to application data. Hadoop YARN: A framework for job scheduling and cluster resource management. Hadoop MapReduce: A YARN-based system for parallel processing of large data sets. http://hadoop.apache.org 2 E6893 Big Data Analytics – Lecture 2: Big Data Platform © 2023 CY Lin, Columbia University Remind -- Hadoop-related Apache Projects 3 Ambari™: A web-based tool for provisioning, managing, and monitoring Hadoop clusters.It also provides a dashboard for viewing cluster health and ability to view MapReduce, Pig and Hive applications visually. Avro™: A data serialization system. Cassandra™: A scalable multi-master database with no single points of failure. Chukwa™: A data collection system for managing large distributed systems. HBase™: A scalable, distributed database that sup
자주 묻는 질문
이 문서는 무료인가요?
네. “Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)” 문서는 무료입니다. 로그인 후 '다운로드'를 클릭하여 원본 파일을 받으세요.
이 문서는 몇 페이지로 되어 있나요?
이 문서는 79페이지입니다, Big Data 과정용. 다운로드하기 전에 온라인으로 미리 볼 수 있습니다.
다운로드하기 전에 미리 볼 수 있나요?
네. 이 페이지의 온라인 리더를 통해 문서를 미리 본 후 다운로드 여부를 결정할 수 있습니다.
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)
미리보기 생성 중...
Trích nội dung tài liệu
E6893 Big Data Analytics Lecture 2: Big Data Analytics Platforms Ching-Yung Lin, Ph.D. Adjunct Professor, Depts. of Electrical Engineering and Computer Science September 15th, 2023 1 E6893 Big Data Analytics – Lecture 2: Big Data Platform © 2023 CY Lin, Columbia University Remind -- Apache Hadoop The Apache™ Hadoop® project develops open-source software for reliable, scalable, distributed computing. The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures. The project includes these modules: Hadoop Common: The common utilities that support the other Hadoop modules. Hadoop Distributed File System (HDFS™): A distributed file system that provides highthroughput access to application data. Hadoop YARN: A framework for job scheduling and cluster resource management. Hadoop MapReduce: A YARN-based system for parallel processing of large data sets. http://hadoop.apache.org 2 E6893 Big Data Analytics – Lecture 2: Big Data Platform © 2023 CY Lin, Columbia University Remind -- Hadoop-related Apache Projects 3 Ambari™: A web-based tool for provisioning, managing, and monitoring Hadoop clusters.It also provides a dashboard for viewing cluster health and ability to view MapReduce, Pig and Hive applications visually. Avro™: A data serialization system. Cassandra™: A scalable multi-master database with no single points of failure. Chukwa™: A data collection system for managing large distributed systems. HBase™: A scalable, distributed database that sup
- 문서명
- Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)
- 학교 / 강의
- Columbia University · Big Data
- 작성자 (문서 내)
- Ching-Yung Lin, Ph.D.
- 내용
- Giới thiệu về Apache Hadoop và các dự án liên quan
- 목차
- 이 문서는 명확한 목차가 없습니다.
- 페이지 수
- 79 페이지
- 업로더
- Uni24h
댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!
Stream (11) (Xử lý luồng dữ liệu) - Julian M. Kunkel
Krone (09) (Sự phát triển của dữ liệu) (Tiếng Anh)
Parallel mf (09) (Thuật toán phân tán phân tích ma trận dữ liệu lớn)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 5)
NoSQL db (06) (Cơ sở dữ liệu NoSQL)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!