Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)
- 页数
- 55
- 格式
- 大小
- 5.4 MB
- 年份
- 2009
- 浏览量
- 0
- 评论
- 0
- 下载次数
- 0
Tài liệu slide bài giảng về Big Data Analytics, giới thiệu các khái niệm cơ bản, MapReduce, ví dụ WordCount, kiến trúc HDFS và Apache Spark.
常见问题
此文档免费吗?
是的。“Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)”是免费的 — 只需登录并点击“下载”即可获取原始文件。
这份文档有多少页?
该文档共有 55 页。您可以在下载前进行在线预览。
我可以在下载前预览吗?
是的。您可以通过在线阅读器直接在本页面预览此文档,然后再决定是否下载。
- 文档名称
- Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)
- 内容
- Tài liệu này cung cấp cái nhìn tổng quan về Big Data Analytics, giải thích các khái niệm cốt lõi và giới thiệu các công nghệ quan trọng như Hadoop và Spark. Nó mô tả cách thức hoạt động của MapReduce và kiến trúc HDFS, đồng thời nhấn mạnh khả năng xử lý nhanh chóng của Spark.
- 目录
- What is Big Data?
- How to analyze Big Data?
- Map/Reduce
- Basic Example: Word Count (Spark & Python)
- Basic Example: Word Count (Spark & Scala)
- Map/Reduce – Parallel Computing
- Map/Reduce History
- Amazon Elastic MapReduce
- HDFS Architecture
- Hadoop & Object Storage
- Apache Spark
- 页数
- 55 页
- 上传者
- Uni24h
正在生成预览...
描述
Big Data Analytics Hadoop and Spark Shelly Garion, Ph.D. IBM Research Haifa 1 What is Big Data? 2 What is Big Data? Big data usually includes data sets with sizes beyond the ability of commonly used software tools to capture, curate, manage, and process data within a tolerable elapsed time. Big data "size" is a constantly moving target, as of 2012 ranging from a few dozen terabytes to many petabytes of data. Big data is a set of techniques and technologies that require new forms of integration to uncover large hidden values from large datasets that are diverse, complex, and of a massive scale. (From Wikipedia) 3 What is Big Data? Volume. Value. Variety. Velocity - the speed of generation of data. Variability - the inconsistency which can be shown by the data at times. Veracity - the quality of the data being captured can vary greatly. Complexity - data management can become a very complex process, especially when large volumes of data come from multiple sources. (From Wikipedia) 4 How to analyze Big Data? 5 Map/Reduce 6 Map/Reduce MapReduce is a framework for processing parallelizable problems across huge datasets using a large number of computers (nodes), collectively referred to as a cluster. Map step: Each worker node applies the "map()" function to the local data, and writes the output to a temporary storage. A master node orchestrates that for redundant copies of input data, only one is processed. Shuffle step: Worker nodes redistribute data based on the output keys (produced by the "map()" function), such that all data belonging to one key is located on the same worker node. Reduce step: Worker nodes now process each group of output data, per key, in parallel. (From Wikipedia) 7 Basic Example: Word Count (Spark & Python) 8 Basic Example: Word Count (Spark & Scala) 9 Map/Reduce – Parallel Computing No dependency among data Data can be split into equal-size chunks Each proces
Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)
正在生成预览...
Big Data Analytics Hadoop and Spark Shelly Garion, Ph.D. IBM Research Haifa 1 What is Big Data? 2 What is Big Data? Big data usually includes data sets with sizes beyond the ability of commonly used software tools to capture, curate, manage, and process data within a tolerable elapsed time. Big data "size" is a constantly moving target, as of 2012 ranging from a few dozen terabytes to many petabytes of data. Big data is a set of techniques and technologies that require new forms of integration to uncover large hidden values from large datasets that are diverse, complex, and of a massive scale. (From Wikipedia) 3 What is Big Data? Volume. Value. Variety. Velocity - the speed of generation of data. Variability - the inconsistency which can be shown by the data at times. Veracity - the quality of the data being captured can vary greatly. Complexity - data management can become a very complex process, especially when large volumes of data come from multiple sources. (From Wikipedia) 4 How to analyze Big Data? 5 Map/Reduce 6 Map/Reduce MapReduce is a framework for processing parallelizable problems across huge datasets using a large number of computers (nodes), collectively referred to as a cluster. Map step: Each worker node applies the "map()" function to the local data, and writes the output to a temporary storage. A master node orchestrates that for redundant copies of input data, only one is processed. Shuffle step: Worker nodes redistribute data based on the output keys (produced by the "map()" function), such that all data belonging to one key is located on the same worker node. Reduce step: Worker nodes now process each group of output data, per key, in parallel. (From Wikipedia) 7 Basic Example: Word Count (Spark & Python) 8 Basic Example: Word Count (Spark & Scala) 9 Map/Reduce – Parallel Computing No dependency among data Data can be split into equal-size chunks Each proces
阅读全文
- 文档名称
- Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)
- 内容
- Tài liệu này cung cấp cái nhìn tổng quan về Big Data Analytics, giải thích các khái niệm cốt lõi và giới thiệu các công nghệ quan trọng như Hadoop và Spark. Nó mô tả cách thức hoạt động của MapReduce và kiến trúc HDFS, đồng thời nhấn mạnh khả năng xử lý nhanh chóng của Spark.
- 目录
- What is Big Data?
- How to analyze Big Data?
- Map/Reduce
- Basic Example: Word Count (Spark & Python)
- Basic Example: Word Count (Spark & Scala)
- Map/Reduce – Parallel Computing
- Map/Reduce History
- Amazon Elastic MapReduce
- HDFS Architecture
- Hadoop & Object Storage
- Apache Spark
- 页数
- 55 页
- 上传者
- Uni24h
评论 (0)
暂无评论。快来抢沙发吧!
Ngân hàng đề thi môn: Hệ thống thông tin quản lý
Đề thi môn Cơ sở dữ liệu (kèm Đáp án) - Đại học Sư phạm kỹ thuật
Đề thi và đáp án môn Hệ thống thông tin kế toán
Đề thi và đáp án môn Cấu trúc dữ liệu giải thuật
Đáp án đề thi môn Mạng máy tính - ĐH Công nghệ thông tin (CNTT)
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
评论 (0)
暂无评论。快来抢沙发吧!