Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
Generating preview...
Slide bài giảng trình bày về xử lý dữ liệu quan hệ trên MapReduce, so sánh với hệ quản trị cơ sở dữ liệu quan hệ, kiến trúc Hadoop YARN và khối lượng công việc OLTP/OLAP.
Description
Fall 2019 Relational Data Processing on MapReduce Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 Peta-scale Data Analysis 30 billion RFID 12+ TBs of tweet data every day tags today (1.3B in 2005) Fall 2019 4.6 billion camera phones world wide ? TBs of data every day 100s of millions of GPS enabled devices sold annually 25+ TBs of log data every day generated by a new user being added every sec. for 3 years 4 billion views/day YouTube is the 2nd most used search engine next to Google 2+ billion 76 million smart meters in 2009… 200M by 2014 people on the Web by end 2011 © 2014 IBM Corporation 2 1 1 Big Data Analysis Fall 2019 A lot of these datasets have some structure Query logs Point-of-sale records User data (e.g., demographics) … How do we perform data analysis at scale? Relational databases and SQL MapReduce (Hadoop) 3 Fall 2019 Relational Databases vs. MapReduce Relational databases: Multipurpose: analysis and transactions; batch and interactive Data integrity via ACID transactions Lots of tools in software ecosystem (for ingesting, reporting, etc.) Supports SQL (and SQL integration, e.g., JDBC) Automatic SQL query optimization MapReduce (Hadoop): Designed for large clusters, fault tolerant Data is accessed in “native format” Supports many query languages Programmers retain control over performance 4 2 2 Fall 2019 Parallel Computation & Data Size Matters! 5 Parallel Relational Databases vs. MapReduce Parallel relational databases Fall 2019 Shared-nothing architecture for parallel processing Schema on “write” Failures are relatively infrequent “Possessive” of data Mostly proprietary MapReduce Schema on “read” Failures are relatively common In situ data processing Open source Hadoop NextGen (YARN) architecture 6 3 3 Fall 2019 MapReduce: A Major Step Backwards? MapReduce is a step backward in database access Separation of the schema
AI summary
- Document name
- Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
- School / Course
- University of Crete · Khai phá dữ liệu
- Content
- Tài liệu so sánh cơ sở dữ liệu quan hệ song song với MapReduce trong việc xử lý dữ liệu lớn, phân tích các ưu nhược điểm và đề xuất giải pháp tích hợp khối lượng công việc OLTP và OLAP.
- Table of contents
- Relational Data Processing on MapReduce
- Peta-scale Data Analysis
- Big Data Analysis
- Relational Databases vs. MapReduce
- Parallel Computation & Data Size Matters!
- Parallel Relational Databases vs. MapReduce
- MapReduce: A Major Step Backwards?
- Map Reduce vs Parallel DBMS
- Database Workloads
- One Database or Two?
- OLTP/OLAP Integration
- Pages
- 34 pages
- Uploaded by
- Uni24h
Frequently asked questions
Is this document free?
Yes. “Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)” is free — just sign in and click Download to get the original file.
How many pages is this document?
The document has 34 pages, for the course Khai phá dữ liệu. You can preview it online before downloading.
Can I preview before downloading?
Yes. You can preview this document right on this page with the online reader, then decide whether to download.
Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
Generating preview...
Fall 2019 Relational Data Processing on MapReduce Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 Peta-scale Data Analysis 30 billion RFID 12+ TBs of tweet data every day tags today (1.3B in 2005) Fall 2019 4.6 billion camera phones world wide ? TBs of data every day 100s of millions of GPS enabled devices sold annually 25+ TBs of log data every day generated by a new user being added every sec. for 3 years 4 billion views/day YouTube is the 2nd most used search engine next to Google 2+ billion 76 million smart meters in 2009… 200M by 2014 people on the Web by end 2011 © 2014 IBM Corporation 2 1 1 Big Data Analysis Fall 2019 A lot of these datasets have some structure Query logs Point-of-sale records User data (e.g., demographics) … How do we perform data analysis at scale? Relational databases and SQL MapReduce (Hadoop) 3 Fall 2019 Relational Databases vs. MapReduce Relational databases: Multipurpose: analysis and transactions; batch and interactive Data integrity via ACID transactions Lots of tools in software ecosystem (for ingesting, reporting, etc.) Supports SQL (and SQL integration, e.g., JDBC) Automatic SQL query optimization MapReduce (Hadoop): Designed for large clusters, fault tolerant Data is accessed in “native format” Supports many query languages Programmers retain control over performance 4 2 2 Fall 2019 Parallel Computation & Data Size Matters! 5 Parallel Relational Databases vs. MapReduce Parallel relational databases Fall 2019 Shared-nothing architecture for parallel processing Schema on “write” Failures are relatively infrequent “Possessive” of data Mostly proprietary MapReduce Schema on “read” Failures are relatively common In situ data processing Open source Hadoop NextGen (YARN) architecture 6 3 3 Fall 2019 MapReduce: A Major Step Backwards? MapReduce is a step backward in database access Separation of the schema
Read full document
- Document name
- Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
- School / Course
- University of Crete · Khai phá dữ liệu
- Content
- Tài liệu so sánh cơ sở dữ liệu quan hệ song song với MapReduce trong việc xử lý dữ liệu lớn, phân tích các ưu nhược điểm và đề xuất giải pháp tích hợp khối lượng công việc OLTP và OLAP.
- Table of contents
- Relational Data Processing on MapReduce
- Peta-scale Data Analysis
- Big Data Analysis
- Relational Databases vs. MapReduce
- Parallel Computation & Data Size Matters!
- Parallel Relational Databases vs. MapReduce
- MapReduce: A Major Step Backwards?
- Map Reduce vs Parallel DBMS
- Database Workloads
- One Database or Two?
- OLTP/OLAP Integration
- Pages
- 34 pages
- Uploaded by
- Uni24h
Comments (0)
No comments yet. Be the first!
Frequent Item Sets Association Rules (Tập phổ biến và Luật kết hợp)
Entity Resolution in the Web of Data (Phân giải thực thể trong Web dữ liệu)
IoT Data Analytics (Phân tích dữ liệu trong Internet vạn vật)
Big Data Processing and Analytics Intro (Xử lý và phân tích dữ liệu lớn)
Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Comments (0)
No comments yet. Be the first!