Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)
- 페이지 수
- 55
- 형식
- 크기
- 5.4 MB
- 연도
- 2009
- 조회수
- 0
- 댓글
- 0
- Lượt tải
- 0
미리보기 생성 중...
Tài liệu slide bài giảng về Big Data Analytics, giới thiệu các khái niệm cơ bản, MapReduce, ví dụ WordCount, kiến trúc HDFS và Apache Spark.
- 문서명
- Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)
- 내용
- Tài liệu này cung cấp cái nhìn tổng quan về Big Data Analytics, giải thích các khái niệm cốt lõi và giới thiệu các công nghệ quan trọng như Hadoop và Spark. Nó mô tả cách thức hoạt động của MapReduce và kiến trúc HDFS, đồng thời nhấn mạnh khả năng xử lý nhanh chóng của Spark.
- 목차
- What is Big Data?
- How to analyze Big Data?
- Map/Reduce
- Basic Example: Word Count (Spark & Python)
- Basic Example: Word Count (Spark & Scala)
- Map/Reduce – Parallel Computing
- Map/Reduce History
- Amazon Elastic MapReduce
- HDFS Architecture
- Hadoop & Object Storage
- Apache Spark
- 페이지 수
- 55 페이지
- 업로더
- Uni24h
설명
Trích nội dung tài liệu
Big Data Analytics Hadoop and Spark Shelly Garion, Ph.D. IBM Research Haifa 1 What is Big Data? 2 What is Big Data? Big data usually includes data sets with sizes beyond the ability of commonly used software tools to capture, curate, manage, and process data within a tolerable elapsed time. Big data "size" is a constantly moving target, as of 2012 ranging from a few dozen terabytes to many petabytes of data. Big data is a set of techniques and technologies that require new forms of integration to uncover large hidden values from large datasets that are diverse, complex, and of a massive scale. (From Wikipedia) 3 What is Big Data? Volume. Value. Variety. Velocity - the speed of generation of data. Variability - the inconsistency which can be shown by the data at times. Veracity - the quality of the data being captured can vary greatly. Complexity - data management can become a very complex process, especially when large volumes of data come from multiple sources. (From Wikipedia) 4 How to analyze Big Data? 5 Map/Reduce 6 Map/Reduce MapReduce is a framework for processing parallelizable problems across huge datasets using a large number of computers (nodes), collectively referred to as a cluster. Map step: Each worker node applies the "map()" function to the local data, and writes the output to a temporary storage. A master node orchestrates that for redundant copies of input data, only one is processed. Shuffle step: Worker nodes redistribute data based on the output keys (produced by the "map()" function), such that all data belonging to one key is located on the same worker node. Reduce step: Worker nodes now process each group of output data, per key, in parallel. (From Wikipedia) 7 Basic Example: Word Count (Spark & Python) 8 Basic Example: Word Count (Spark & Scala) 9 Map/Reduce – Parallel Computing No dependency among data Data can be split into equal-size chunks Each proces
자주 묻는 질문
이 문서는 무료인가요?
네. “Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)” 문서는 무료입니다. 로그인 후 '다운로드'를 클릭하여 원본 파일을 받으세요.
이 문서는 몇 페이지로 되어 있나요?
이 문서는 55페이지입니다. 다운로드하기 전에 온라인으로 미리 볼 수 있습니다.
다운로드하기 전에 미리 볼 수 있나요?
네. 이 페이지의 온라인 리더를 통해 문서를 미리 본 후 다운로드 여부를 결정할 수 있습니다.
Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)
미리보기 생성 중...
Trích nội dung tài liệu
Big Data Analytics Hadoop and Spark Shelly Garion, Ph.D. IBM Research Haifa 1 What is Big Data? 2 What is Big Data? Big data usually includes data sets with sizes beyond the ability of commonly used software tools to capture, curate, manage, and process data within a tolerable elapsed time. Big data "size" is a constantly moving target, as of 2012 ranging from a few dozen terabytes to many petabytes of data. Big data is a set of techniques and technologies that require new forms of integration to uncover large hidden values from large datasets that are diverse, complex, and of a massive scale. (From Wikipedia) 3 What is Big Data? Volume. Value. Variety. Velocity - the speed of generation of data. Variability - the inconsistency which can be shown by the data at times. Veracity - the quality of the data being captured can vary greatly. Complexity - data management can become a very complex process, especially when large volumes of data come from multiple sources. (From Wikipedia) 4 How to analyze Big Data? 5 Map/Reduce 6 Map/Reduce MapReduce is a framework for processing parallelizable problems across huge datasets using a large number of computers (nodes), collectively referred to as a cluster. Map step: Each worker node applies the "map()" function to the local data, and writes the output to a temporary storage. A master node orchestrates that for redundant copies of input data, only one is processed. Shuffle step: Worker nodes redistribute data based on the output keys (produced by the "map()" function), such that all data belonging to one key is located on the same worker node. Reduce step: Worker nodes now process each group of output data, per key, in parallel. (From Wikipedia) 7 Basic Example: Word Count (Spark & Python) 8 Basic Example: Word Count (Spark & Scala) 9 Map/Reduce – Parallel Computing No dependency among data Data can be split into equal-size chunks Each proces
- 문서명
- Big Data Analytics Hadoop and Spark (Phân tích dữ liệu lớn Hadoop và Spark) (Tiếng Anh)
- 내용
- Tài liệu này cung cấp cái nhìn tổng quan về Big Data Analytics, giải thích các khái niệm cốt lõi và giới thiệu các công nghệ quan trọng như Hadoop và Spark. Nó mô tả cách thức hoạt động của MapReduce và kiến trúc HDFS, đồng thời nhấn mạnh khả năng xử lý nhanh chóng của Spark.
- 목차
- What is Big Data?
- How to analyze Big Data?
- Map/Reduce
- Basic Example: Word Count (Spark & Python)
- Basic Example: Word Count (Spark & Scala)
- Map/Reduce – Parallel Computing
- Map/Reduce History
- Amazon Elastic MapReduce
- HDFS Architecture
- Hadoop & Object Storage
- Apache Spark
- 페이지 수
- 55 페이지
- 업로더
- Uni24h
댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!
Ngân hàng đề thi môn: Hệ thống thông tin quản lý
Đề thi môn Cơ sở dữ liệu (kèm Đáp án) - Đại học Sư phạm kỹ thuật
Đề thi và đáp án môn Hệ thống thông tin kế toán
Đề thi và đáp án môn Cấu trúc dữ liệu giải thuật
Đáp án đề thi môn Mạng máy tính - ĐH Công nghệ thông tin (CNTT)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!