Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
Generating preview...
Tài liệu slide bài giảng giới thiệu về phân tích dữ liệu quy mô lớn sử dụng Apache Spark, bao gồm các vấn đề dữ liệu lớn, MapReduce, và Spark.
Description
Fall 2019 Introduction to Scalable Data Analytics using Apache Spark Fall 2019 Outline ⬤ Big Data Problems: Distributing Work, Failures, Slow Machines ⬤ What is Apache Spark? ⬤ Core things of Apache Spark ⬥ RDD ⬤ Core Functionality of Apache Spark ⬤ Simple tutorial Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 2 Fall 2019 Fall 2019 Hardware for Big Data Big Data Problems: Distributing Work, Failures, Slow Machines Bunch of Hard Drives …. and CPUs The Big Data Problem ⬥ Data growing faster than CPU speeds ⬥ Data growing faster than per-machine storage ⬤ Can’t process or store all data on one machine ⬤ 3 4 Fall 2019 Fall 2019 Hardware for Big Data ⬤ One big box ! (1990s solution) ⬥ All processors share memory ⬤ Very expensive ⬥ Low volume ⬥ All “premium” HW ⬤ Still not big enough! Hardware for Big Data Image: Wikimedia Commons / User:Tonusamuel ⬤ Consumer-grade hardware ⬥ Not "gold plated" ⬤ Many desktop-like servers ⬥ Easy to add capacity ⬥ Cheaper per CPU/disk ⬤ But, implies complexity in software Image: Steve Jurvetson/Flickr 5 6 Fall 2019 Fall 2019 Problems with Cheap HW ⬤ ⬤ ⬤ The Opportunity Failures, e.g. (Google numbers) ⬥ 1-5% hard drives/year ⬥ 0.2% DIMMs/year ⬤ Cluster computing is a game-changer! ⬤ Provides access to low-cost computing and storage Network speeds vs. shared memory ⬥ Much more latency ⬥ Network slower than storage ⬤ Costs decreasing every year ⬤ The challenge is programming the resources Uneven performance ⬤ What’s hard about Cluster computing? ⬥ How do we split work across machines? Google Datacenter 7 8 Fall 2019 How do you Count the Number of Occurrences of each Word in a Document? “I am Sam I am Sam Sam I am Do you like Green eggs and ham?” Fall 2019 Centrilized Approach: Use a Hash Table I: 3 am: 3 Sam: 3 do: 1 you: 1 like: 1 … “I am Sam I am Sam Sam I am Do you like Green eggs and ham?” 9 {} 10 Fall 2019 Centri
AI summary
- Document name
- Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
- School / Course
- University of Crete · Khai phá dữ liệu
- Content
- Tài liệu giới thiệu Apache Spark cho phân tích dữ liệu lớn, nêu bật các vấn đề của Big Data và cách Spark giải quyết chúng. Nó bao gồm các khái niệm cốt lõi như RDD và so sánh với MapReduce.
- Table of contents
- Big Data Problems: Distributing Work, Failures, Slow Machines
- What is Apache Spark?
- Core things of Apache Spark
- Core Functionality of Apache Spark
- Simple tutorial
- Pages
- 43 pages
- Uploaded by
- Uni24h
Frequently asked questions
Is this document free?
Yes. “Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)” is free — just sign in and click Download to get the original file.
How many pages is this document?
The document has 43 pages, for the course Khai phá dữ liệu. You can preview it online before downloading.
Can I preview before downloading?
Yes. You can preview this document right on this page with the online reader, then decide whether to download.
Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
Generating preview...
Fall 2019 Introduction to Scalable Data Analytics using Apache Spark Fall 2019 Outline ⬤ Big Data Problems: Distributing Work, Failures, Slow Machines ⬤ What is Apache Spark? ⬤ Core things of Apache Spark ⬥ RDD ⬤ Core Functionality of Apache Spark ⬤ Simple tutorial Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 2 Fall 2019 Fall 2019 Hardware for Big Data Big Data Problems: Distributing Work, Failures, Slow Machines Bunch of Hard Drives …. and CPUs The Big Data Problem ⬥ Data growing faster than CPU speeds ⬥ Data growing faster than per-machine storage ⬤ Can’t process or store all data on one machine ⬤ 3 4 Fall 2019 Fall 2019 Hardware for Big Data ⬤ One big box ! (1990s solution) ⬥ All processors share memory ⬤ Very expensive ⬥ Low volume ⬥ All “premium” HW ⬤ Still not big enough! Hardware for Big Data Image: Wikimedia Commons / User:Tonusamuel ⬤ Consumer-grade hardware ⬥ Not "gold plated" ⬤ Many desktop-like servers ⬥ Easy to add capacity ⬥ Cheaper per CPU/disk ⬤ But, implies complexity in software Image: Steve Jurvetson/Flickr 5 6 Fall 2019 Fall 2019 Problems with Cheap HW ⬤ ⬤ ⬤ The Opportunity Failures, e.g. (Google numbers) ⬥ 1-5% hard drives/year ⬥ 0.2% DIMMs/year ⬤ Cluster computing is a game-changer! ⬤ Provides access to low-cost computing and storage Network speeds vs. shared memory ⬥ Much more latency ⬥ Network slower than storage ⬤ Costs decreasing every year ⬤ The challenge is programming the resources Uneven performance ⬤ What’s hard about Cluster computing? ⬥ How do we split work across machines? Google Datacenter 7 8 Fall 2019 How do you Count the Number of Occurrences of each Word in a Document? “I am Sam I am Sam Sam I am Do you like Green eggs and ham?” Fall 2019 Centrilized Approach: Use a Hash Table I: 3 am: 3 Sam: 3 do: 1 you: 1 like: 1 … “I am Sam I am Sam Sam I am Do you like Green eggs and ham?” 9 {} 10 Fall 2019 Centri
Read full document
- Document name
- Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
- School / Course
- University of Crete · Khai phá dữ liệu
- Content
- Tài liệu giới thiệu Apache Spark cho phân tích dữ liệu lớn, nêu bật các vấn đề của Big Data và cách Spark giải quyết chúng. Nó bao gồm các khái niệm cốt lõi như RDD và so sánh với MapReduce.
- Table of contents
- Big Data Problems: Distributing Work, Failures, Slow Machines
- What is Apache Spark?
- Core things of Apache Spark
- Core Functionality of Apache Spark
- Simple tutorial
- Pages
- 43 pages
- Uploaded by
- Uni24h
Comments (0)
No comments yet. Be the first!
Basic association analysis (Chap 6) (Thuật toán trong phân tích kết hợp)
Frequent Item Sets Association Rules (Tập phổ biến và Luật kết hợp)
OLAP, Data Warehouse, and Column Store (Lecture 3) (Khái niệm về cơ sở dữ liệu quan hệ, OLTP, OLAP và kiến trúc kho dữ liệu)
Map Reduce Algorithm Design (Lecture 6) (Thiết kế thuật toán MapReduce)
Hadoop Programming (Khái niệm cơ bản về lập trình Hadoop MapReduce)
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Comments (0)
No comments yet. Be the first!