Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
- 페이지 수
- 43
- 형식
- 크기
- 7.5 MB
- 연도
- 2019
- Trường
- University of Crete
- 조회수
- 0
- 댓글
- 0
- Lượt tải
- 0
미리보기 생성 중...
Tài liệu slide bài giảng giới thiệu về phân tích dữ liệu quy mô lớn sử dụng Apache Spark, bao gồm các vấn đề dữ liệu lớn, MapReduce, và Spark.
- 문서명
- Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
- 학교 / 강의
- University of Crete · Khai phá dữ liệu
- 내용
- Tài liệu giới thiệu Apache Spark cho phân tích dữ liệu lớn, nêu bật các vấn đề của Big Data và cách Spark giải quyết chúng. Nó bao gồm các khái niệm cốt lõi như RDD và so sánh với MapReduce.
- 목차
- Big Data Problems: Distributing Work, Failures, Slow Machines
- What is Apache Spark?
- Core things of Apache Spark
- Core Functionality of Apache Spark
- Simple tutorial
- 페이지 수
- 43 페이지
- 업로더
- Uni24h
설명
Trích nội dung tài liệu
Fall 2019 Introduction to Scalable Data Analytics using Apache Spark Fall 2019 Outline ⬤ Big Data Problems: Distributing Work, Failures, Slow Machines ⬤ What is Apache Spark? ⬤ Core things of Apache Spark ⬥ RDD ⬤ Core Functionality of Apache Spark ⬤ Simple tutorial Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 2 Fall 2019 Fall 2019 Hardware for Big Data Big Data Problems: Distributing Work, Failures, Slow Machines Bunch of Hard Drives …. and CPUs The Big Data Problem ⬥ Data growing faster than CPU speeds ⬥ Data growing faster than per-machine storage ⬤ Can’t process or store all data on one machine ⬤ 3 4 Fall 2019 Fall 2019 Hardware for Big Data ⬤ One big box ! (1990s solution) ⬥ All processors share memory ⬤ Very expensive ⬥ Low volume ⬥ All “premium” HW ⬤ Still not big enough! Hardware for Big Data Image: Wikimedia Commons / User:Tonusamuel ⬤ Consumer-grade hardware ⬥ Not "gold plated" ⬤ Many desktop-like servers ⬥ Easy to add capacity ⬥ Cheaper per CPU/disk ⬤ But, implies complexity in software Image: Steve Jurvetson/Flickr 5 6 Fall 2019 Fall 2019 Problems with Cheap HW ⬤ ⬤ ⬤ The Opportunity Failures, e.g. (Google numbers) ⬥ 1-5% hard drives/year ⬥ 0.2% DIMMs/year ⬤ Cluster computing is a game-changer! ⬤ Provides access to low-cost computing and storage Network speeds vs. shared memory ⬥ Much more latency ⬥ Network slower than storage ⬤ Costs decreasing every year ⬤ The challenge is programming the resources Uneven performance ⬤ What’s hard about Cluster computing? ⬥ How do we split work across machines? Google Datacenter 7 8 Fall 2019 How do you Count the Number of Occurrences of each Word in a Document? “I am Sam I am Sam Sam I am Do you like Green eggs and ham?” Fall 2019 Centrilized Approach: Use a Hash Table I: 3 am: 3 Sam: 3 do: 1 you: 1 like: 1 … “I am Sam I am Sam Sam I am Do you like Green eggs and ham?” 9 {} 10 Fall 2019 Centri
자주 묻는 질문
이 문서는 무료인가요?
네. “Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)” 문서는 무료입니다. 로그인 후 '다운로드'를 클릭하여 원본 파일을 받으세요.
이 문서는 몇 페이지로 되어 있나요?
이 문서는 43페이지입니다, Khai phá dữ liệu 과정용. 다운로드하기 전에 온라인으로 미리 볼 수 있습니다.
다운로드하기 전에 미리 볼 수 있나요?
네. 이 페이지의 온라인 리더를 통해 문서를 미리 본 후 다운로드 여부를 결정할 수 있습니다.
Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
미리보기 생성 중...
Trích nội dung tài liệu
Fall 2019 Introduction to Scalable Data Analytics using Apache Spark Fall 2019 Outline ⬤ Big Data Problems: Distributing Work, Failures, Slow Machines ⬤ What is Apache Spark? ⬤ Core things of Apache Spark ⬥ RDD ⬤ Core Functionality of Apache Spark ⬤ Simple tutorial Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 2 Fall 2019 Fall 2019 Hardware for Big Data Big Data Problems: Distributing Work, Failures, Slow Machines Bunch of Hard Drives …. and CPUs The Big Data Problem ⬥ Data growing faster than CPU speeds ⬥ Data growing faster than per-machine storage ⬤ Can’t process or store all data on one machine ⬤ 3 4 Fall 2019 Fall 2019 Hardware for Big Data ⬤ One big box ! (1990s solution) ⬥ All processors share memory ⬤ Very expensive ⬥ Low volume ⬥ All “premium” HW ⬤ Still not big enough! Hardware for Big Data Image: Wikimedia Commons / User:Tonusamuel ⬤ Consumer-grade hardware ⬥ Not "gold plated" ⬤ Many desktop-like servers ⬥ Easy to add capacity ⬥ Cheaper per CPU/disk ⬤ But, implies complexity in software Image: Steve Jurvetson/Flickr 5 6 Fall 2019 Fall 2019 Problems with Cheap HW ⬤ ⬤ ⬤ The Opportunity Failures, e.g. (Google numbers) ⬥ 1-5% hard drives/year ⬥ 0.2% DIMMs/year ⬤ Cluster computing is a game-changer! ⬤ Provides access to low-cost computing and storage Network speeds vs. shared memory ⬥ Much more latency ⬥ Network slower than storage ⬤ Costs decreasing every year ⬤ The challenge is programming the resources Uneven performance ⬤ What’s hard about Cluster computing? ⬥ How do we split work across machines? Google Datacenter 7 8 Fall 2019 How do you Count the Number of Occurrences of each Word in a Document? “I am Sam I am Sam Sam I am Do you like Green eggs and ham?” Fall 2019 Centrilized Approach: Use a Hash Table I: 3 am: 3 Sam: 3 do: 1 you: 1 like: 1 … “I am Sam I am Sam Sam I am Do you like Green eggs and ham?” 9 {} 10 Fall 2019 Centri
- 문서명
- Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
- 학교 / 강의
- University of Crete · Khai phá dữ liệu
- 내용
- Tài liệu giới thiệu Apache Spark cho phân tích dữ liệu lớn, nêu bật các vấn đề của Big Data và cách Spark giải quyết chúng. Nó bao gồm các khái niệm cốt lõi như RDD và so sánh với MapReduce.
- 목차
- Big Data Problems: Distributing Work, Failures, Slow Machines
- What is Apache Spark?
- Core things of Apache Spark
- Core Functionality of Apache Spark
- Simple tutorial
- 페이지 수
- 43 페이지
- 업로더
- Uni24h
댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!
Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
IoT Data Analytics (Phân tích dữ liệu trong Internet vạn vật)
Colocation 2 (Khám phá mẫu colocation trong dữ liệu không gian) - Zhe Jiang
Big Data Processing and Analytics Intro (Xử lý và phân tích dữ liệu lớn)
Basic association analysis (Chap 6) (Thuật toán trong phân tích kết hợp)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!