Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey
- 页数
- 513
- 格式
- 大小
- 2.9 MB
- 语言
- EN
- 年份
- 2014
- 浏览量
- 0
- 评论
- 0
- 下载次数
- 0
Cuốn sách này tập trung vào khai phá dữ liệu lớn, bao gồm các thuật toán phân tán, tìm kiếm tương tự, xử lý luồng dữ liệu, công nghệ tìm kiếm, khai phá tập phổ biến, phân cụm, quản lý quảng cáo, hệ thống gợi ý, phân tích đồ thị lớn, giảm chiều dữ liệu và các thuật toán học máy cho dữ liệu lớn.
常见问题
此文档免费吗?
是的。“Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey”是免费的 — 只需登录并点击“下载”即可获取原始文件。
这份文档有多少页?
该文档共有 513 页。您可以在下载前进行在线预览。
我可以在下载前预览吗?
是的。您可以通过在线阅读器直接在本页面预览此文档,然后再决定是否下载。
- 文档名称
- Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey
- 作者(文档中)
- Anand Rajaraman, Jure Leskovec, Jeffrey D. Ullman
- 内容
- Đây là một cuốn sách về khai thác dữ liệu lớn, tập trung vào các thuật toán xử lý dữ liệu không vừa bộ nhớ chính. Sách bao gồm các chủ đề từ hệ thống phân tán, tìm kiếm tương tự, xử lý luồng, công cụ tìm kiếm, đến phân cụm, đồ thị lớn và học máy cho dữ liệu quy mô lớn.
- 目录
- 2 MapReduce and the New Software Stack
- 2.1 Distributed File Systems
- 2.1.1 Physical Organization of Compute Nodes
- 2.1.2 Large-Scale File-System Organization
- 2.2 MapReduce
- 2.2.1 The Map Tasks
- 2.2.2 Grouping by Key
- 2.2.3 The Reduce Tasks
- 2.2.4 Combiners
- 页数
- 513 页
- 上传者
- Uni24h
正在生成预览...
描述
Mining of Massive Datasets Jure Leskovec Stanford Univ. Anand Rajaraman Milliway Labs Jeffrey D. Ullman Stanford Univ. Copyright c 2010, 2011, 2012, 2013, 2014 Anand Rajaraman, Jure Leskovec, and Jeffrey D. Ullman ii Preface This book evolved from material developed over several years by Anand Rajaraman and Jeff Ullman for a one-quarter course at Stanford. The course CS345A, titled “Web Mining,” was designed as an advanced graduate course, although it has become accessible and interesting to advanced undergraduates. When Jure Leskovec joined the Stanford faculty, we reorganized the material considerably. He introduced a new course CS224W on network analysis and added material to CS345A, which was renumbered CS246. The three authors also introduced a large-scale data-mining project course, CS341. The book now contains material taught in all three courses. What the Book Is About At the highest level of description, this book is about data mining. However, it focuses on data mining of very large amounts of data, that is, data so large it does not fit in main memory. Because of the emphasis on size, many of our examples are about the Web or data derived from the Web. Further, the book takes an algorithmic point of view: data mining is about applying algorithms to data, rather than using data to “train” a machine-learning engine of some sort. The principal topics covered are: 1. Distributed file systems and map-reduce as a tool for creating parallel algorithms that succeed on very large amounts of data. 2. Similarity search, including the key techniques of minhashing and localitysensitive hashing. 3. Data-stream processing and specialized algorithms for dealing with data that arrives so fast it must be processed immediately or lost. 4. The technology of search engines, including Google’s PageRank, link-spam detection, and the hubs-and-authorities approach. 5. Frequent-itemset mining, including association rules, market-baskets, the A-Priori Algorithm and its impr
Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey
正在生成预览...
Mining of Massive Datasets Jure Leskovec Stanford Univ. Anand Rajaraman Milliway Labs Jeffrey D. Ullman Stanford Univ. Copyright c 2010, 2011, 2012, 2013, 2014 Anand Rajaraman, Jure Leskovec, and Jeffrey D. Ullman ii Preface This book evolved from material developed over several years by Anand Rajaraman and Jeff Ullman for a one-quarter course at Stanford. The course CS345A, titled “Web Mining,” was designed as an advanced graduate course, although it has become accessible and interesting to advanced undergraduates. When Jure Leskovec joined the Stanford faculty, we reorganized the material considerably. He introduced a new course CS224W on network analysis and added material to CS345A, which was renumbered CS246. The three authors also introduced a large-scale data-mining project course, CS341. The book now contains material taught in all three courses. What the Book Is About At the highest level of description, this book is about data mining. However, it focuses on data mining of very large amounts of data, that is, data so large it does not fit in main memory. Because of the emphasis on size, many of our examples are about the Web or data derived from the Web. Further, the book takes an algorithmic point of view: data mining is about applying algorithms to data, rather than using data to “train” a machine-learning engine of some sort. The principal topics covered are: 1. Distributed file systems and map-reduce as a tool for creating parallel algorithms that succeed on very large amounts of data. 2. Similarity search, including the key techniques of minhashing and localitysensitive hashing. 3. Data-stream processing and specialized algorithms for dealing with data that arrives so fast it must be processed immediately or lost. 4. The technology of search engines, including Google’s PageRank, link-spam detection, and the hubs-and-authorities approach. 5. Frequent-itemset mining, including association rules, market-baskets, the A-Priori Algorithm and its impr
阅读全文
- 文档名称
- Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey
- 作者(文档中)
- Anand Rajaraman, Jure Leskovec, Jeffrey D. Ullman
- 内容
- Đây là một cuốn sách về khai thác dữ liệu lớn, tập trung vào các thuật toán xử lý dữ liệu không vừa bộ nhớ chính. Sách bao gồm các chủ đề từ hệ thống phân tán, tìm kiếm tương tự, xử lý luồng, công cụ tìm kiếm, đến phân cụm, đồ thị lớn và học máy cho dữ liệu quy mô lớn.
- 目录
- 2 MapReduce and the New Software Stack
- 2.1 Distributed File Systems
- 2.1.1 Physical Organization of Compute Nodes
- 2.1.2 Large-Scale File-System Organization
- 2.2 MapReduce
- 2.2.1 The Map Tasks
- 2.2.2 Grouping by Key
- 2.2.3 The Reduce Tasks
- 2.2.4 Combiners
- 页数
- 513 页
- 上传者
- Uni24h
评论 (0)
暂无评论。快来抢沙发吧!
Ngân hàng đề thi môn: Hệ thống thông tin quản lý
Đề thi môn Cơ sở dữ liệu (kèm Đáp án) - Đại học Sư phạm kỹ thuật
Đề thi và đáp án môn Hệ thống thông tin kế toán
Đề thi và đáp án môn Cấu trúc dữ liệu giải thuật
Đáp án đề thi môn Mạng máy tính - ĐH Công nghệ thông tin (CNTT)
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

评论 (0)
暂无评论。快来抢沙发吧!