Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
正在生成预览...
Slide bài giảng trình bày về xử lý dữ liệu quan hệ trên MapReduce, so sánh với hệ quản trị cơ sở dữ liệu quan hệ, kiến trúc Hadoop YARN và khối lượng công việc OLTP/OLAP.
描述
Fall 2019 Relational Data Processing on MapReduce Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 Peta-scale Data Analysis 30 billion RFID 12+ TBs of tweet data every day tags today (1.3B in 2005) Fall 2019 4.6 billion camera phones world wide ? TBs of data every day 100s of millions of GPS enabled devices sold annually 25+ TBs of log data every day generated by a new user being added every sec. for 3 years 4 billion views/day YouTube is the 2nd most used search engine next to Google 2+ billion 76 million smart meters in 2009… 200M by 2014 people on the Web by end 2011 © 2014 IBM Corporation 2 1 1 Big Data Analysis Fall 2019 A lot of these datasets have some structure Query logs Point-of-sale records User data (e.g., demographics) … How do we perform data analysis at scale? Relational databases and SQL MapReduce (Hadoop) 3 Fall 2019 Relational Databases vs. MapReduce Relational databases: Multipurpose: analysis and transactions; batch and interactive Data integrity via ACID transactions Lots of tools in software ecosystem (for ingesting, reporting, etc.) Supports SQL (and SQL integration, e.g., JDBC) Automatic SQL query optimization MapReduce (Hadoop): Designed for large clusters, fault tolerant Data is accessed in “native format” Supports many query languages Programmers retain control over performance 4 2 2 Fall 2019 Parallel Computation & Data Size Matters! 5 Parallel Relational Databases vs. MapReduce Parallel relational databases Fall 2019 Shared-nothing architecture for parallel processing Schema on “write” Failures are relatively infrequent “Possessive” of data Mostly proprietary MapReduce Schema on “read” Failures are relatively common In situ data processing Open source Hadoop NextGen (YARN) architecture 6 3 3 Fall 2019 MapReduce: A Major Step Backwards? MapReduce is a step backward in database access Separation of the schema
AI 摘要
- 文档名称
- Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
- 学校 / 课程
- University of Crete · Khai phá dữ liệu
- 内容
- Tài liệu so sánh cơ sở dữ liệu quan hệ song song với MapReduce trong việc xử lý dữ liệu lớn, phân tích các ưu nhược điểm và đề xuất giải pháp tích hợp khối lượng công việc OLTP và OLAP.
- 目录
- Relational Data Processing on MapReduce
- Peta-scale Data Analysis
- Big Data Analysis
- Relational Databases vs. MapReduce
- Parallel Computation & Data Size Matters!
- Parallel Relational Databases vs. MapReduce
- MapReduce: A Major Step Backwards?
- Map Reduce vs Parallel DBMS
- Database Workloads
- One Database or Two?
- OLTP/OLAP Integration
- 页数
- 34 页
- 上传者
- Uni24h
常见问题
此文档免费吗?
是的。“Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)”是免费的 — 只需登录并点击“下载”即可获取原始文件。
这份文档有多少页?
该文档共有 34 页,适用于课程 Khai phá dữ liệu。您可以在下载前进行在线预览。
我可以在下载前预览吗?
是的。您可以通过在线阅读器直接在本页面预览此文档,然后再决定是否下载。
Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
正在生成预览...
Fall 2019 Relational Data Processing on MapReduce Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 Peta-scale Data Analysis 30 billion RFID 12+ TBs of tweet data every day tags today (1.3B in 2005) Fall 2019 4.6 billion camera phones world wide ? TBs of data every day 100s of millions of GPS enabled devices sold annually 25+ TBs of log data every day generated by a new user being added every sec. for 3 years 4 billion views/day YouTube is the 2nd most used search engine next to Google 2+ billion 76 million smart meters in 2009… 200M by 2014 people on the Web by end 2011 © 2014 IBM Corporation 2 1 1 Big Data Analysis Fall 2019 A lot of these datasets have some structure Query logs Point-of-sale records User data (e.g., demographics) … How do we perform data analysis at scale? Relational databases and SQL MapReduce (Hadoop) 3 Fall 2019 Relational Databases vs. MapReduce Relational databases: Multipurpose: analysis and transactions; batch and interactive Data integrity via ACID transactions Lots of tools in software ecosystem (for ingesting, reporting, etc.) Supports SQL (and SQL integration, e.g., JDBC) Automatic SQL query optimization MapReduce (Hadoop): Designed for large clusters, fault tolerant Data is accessed in “native format” Supports many query languages Programmers retain control over performance 4 2 2 Fall 2019 Parallel Computation & Data Size Matters! 5 Parallel Relational Databases vs. MapReduce Parallel relational databases Fall 2019 Shared-nothing architecture for parallel processing Schema on “write” Failures are relatively infrequent “Possessive” of data Mostly proprietary MapReduce Schema on “read” Failures are relatively common In situ data processing Open source Hadoop NextGen (YARN) architecture 6 3 3 Fall 2019 MapReduce: A Major Step Backwards? MapReduce is a step backward in database access Separation of the schema
阅读全文
- 文档名称
- Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
- 学校 / 课程
- University of Crete · Khai phá dữ liệu
- 内容
- Tài liệu so sánh cơ sở dữ liệu quan hệ song song với MapReduce trong việc xử lý dữ liệu lớn, phân tích các ưu nhược điểm và đề xuất giải pháp tích hợp khối lượng công việc OLTP và OLAP.
- 目录
- Relational Data Processing on MapReduce
- Peta-scale Data Analysis
- Big Data Analysis
- Relational Databases vs. MapReduce
- Parallel Computation & Data Size Matters!
- Parallel Relational Databases vs. MapReduce
- MapReduce: A Major Step Backwards?
- Map Reduce vs Parallel DBMS
- Database Workloads
- One Database or Two?
- OLTP/OLAP Integration
- 页数
- 34 页
- 上传者
- Uni24h
评论 (0)
暂无评论。快来抢沙发吧!
Frequent Item Sets Association Rules (Tập phổ biến và Luật kết hợp)
Entity Resolution in the Web of Data (Phân giải thực thể trong Web dữ liệu)
IoT Data Analytics (Phân tích dữ liệu trong Internet vạn vật)
Big Data Processing and Analytics Intro (Xử lý và phân tích dữ liệu lớn)
Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
评论 (0)
暂无评论。快来抢沙发吧!