Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)
Đang tạo bản xem trước...
Bài giảng số 2 về Nền tảng Phân tích Dữ liệu Lớn, tập trung vào Apache Hadoop và các dự án liên quan, cùng các trường hợp sử dụng và ví dụ về MapReduce.
Mô tả
E6893 Big Data Analytics Lecture 2: Big Data Analytics Platforms Ching-Yung Lin, Ph.D. Adjunct Professor, Depts. of Electrical Engineering and Computer Science September 15th, 2023 1 E6893 Big Data Analytics – Lecture 2: Big Data Platform © 2023 CY Lin, Columbia University Remind -- Apache Hadoop The Apache™ Hadoop® project develops open-source software for reliable, scalable, distributed computing. The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures. The project includes these modules: Hadoop Common: The common utilities that support the other Hadoop modules. Hadoop Distributed File System (HDFS™): A distributed file system that provides highthroughput access to application data. Hadoop YARN: A framework for job scheduling and cluster resource management. Hadoop MapReduce: A YARN-based system for parallel processing of large data sets. http://hadoop.apache.org 2 E6893 Big Data Analytics – Lecture 2: Big Data Platform © 2023 CY Lin, Columbia University Remind -- Hadoop-related Apache Projects 3 Ambari™: A web-based tool for provisioning, managing, and monitoring Hadoop clusters.It also provides a dashboard for viewing cluster health and ability to view MapReduce, Pig and Hive applications visually. Avro™: A data serialization system. Cassandra™: A scalable multi-master database with no single points of failure. Chukwa™: A data collection system for managing large distributed systems. HBase™: A scalable, distributed database that sup
Tóm tắt AI
- Tên tài liệu
- Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)
- Trường / Môn
- Columbia University · Big Data
- Tác giả (trong tài liệu)
- Ching-Yung Lin, Ph.D.
- Nội dung
- Giới thiệu về Apache Hadoop và các dự án liên quan
- Mục lục
- Tài liệu không có mục lục rõ ràng.
- Số trang
- 79 trang
- Người đăng
- Uni24h
Câu hỏi thường gặp
Tài liệu này có miễn phí không?
Có. “Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)” miễn phí — bạn chỉ cần đăng nhập rồi bấm Tải xuống để lấy file gốc.
Tài liệu dài bao nhiêu trang?
Tài liệu gồm 79 trang, thuộc môn Big Data. Bạn có thể xem trước online trước khi tải.
Tôi có thể xem trước trước khi tải không?
Có. Bạn xem trước tài liệu ngay trên trang này bằng trình đọc online, rồi quyết định tải về.
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)
Đang tạo bản xem trước...
E6893 Big Data Analytics Lecture 2: Big Data Analytics Platforms Ching-Yung Lin, Ph.D. Adjunct Professor, Depts. of Electrical Engineering and Computer Science September 15th, 2023 1 E6893 Big Data Analytics – Lecture 2: Big Data Platform © 2023 CY Lin, Columbia University Remind -- Apache Hadoop The Apache™ Hadoop® project develops open-source software for reliable, scalable, distributed computing. The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures. The project includes these modules: Hadoop Common: The common utilities that support the other Hadoop modules. Hadoop Distributed File System (HDFS™): A distributed file system that provides highthroughput access to application data. Hadoop YARN: A framework for job scheduling and cluster resource management. Hadoop MapReduce: A YARN-based system for parallel processing of large data sets. http://hadoop.apache.org 2 E6893 Big Data Analytics – Lecture 2: Big Data Platform © 2023 CY Lin, Columbia University Remind -- Hadoop-related Apache Projects 3 Ambari™: A web-based tool for provisioning, managing, and monitoring Hadoop clusters.It also provides a dashboard for viewing cluster health and ability to view MapReduce, Pig and Hive applications visually. Avro™: A data serialization system. Cassandra™: A scalable multi-master database with no single points of failure. Chukwa™: A data collection system for managing large distributed systems. HBase™: A scalable, distributed database that sup
Đọc toàn bộ tài liệu
- Tên tài liệu
- Big Data Analytics - Phân tích dữ liệu lớn (Lecture 2)
- Trường / Môn
- Columbia University · Big Data
- Tác giả (trong tài liệu)
- Ching-Yung Lin, Ph.D.
- Nội dung
- Giới thiệu về Apache Hadoop và các dự án liên quan
- Mục lục
- Tài liệu không có mục lục rõ ràng.
- Số trang
- 79 trang
- Người đăng
- Uni24h
Bình luận (0)
Chưa có bình luận nào. Hãy là người đầu tiên!
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Bình luận (0)
Chưa có bình luận nào. Hãy là người đầu tiên!