Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey
- 페이지 수
- 513
- 형식
- 크기
- 2.9 MB
- 언어
- EN · English
- 연도
- 2014
- 조회수
- 0
- 댓글
- 0
- Lượt tải
- 0
미리보기 생성 중...
Cuốn sách này tập trung vào khai phá dữ liệu lớn, bao gồm các thuật toán phân tán, tìm kiếm tương tự, xử lý luồng dữ liệu, công nghệ tìm kiếm, khai phá tập phổ biến, phân cụm, quản lý quảng cáo, hệ thống gợi ý, phân tích đồ thị lớn, giảm chiều dữ liệu và các thuật toán học máy cho dữ liệu lớn.
- 문서명
- Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey
- 작성자 (문서 내)
- Anand Rajaraman, Jure Leskovec, Jeffrey D. Ullman
- 내용
- Đây là một cuốn sách về khai thác dữ liệu lớn, tập trung vào các thuật toán xử lý dữ liệu không vừa bộ nhớ chính. Sách bao gồm các chủ đề từ hệ thống phân tán, tìm kiếm tương tự, xử lý luồng, công cụ tìm kiếm, đến phân cụm, đồ thị lớn và học máy cho dữ liệu quy mô lớn.
- 목차
- 2 MapReduce and the New Software Stack
- 2.1 Distributed File Systems
- 2.1.1 Physical Organization of Compute Nodes
- 2.1.2 Large-Scale File-System Organization
- 2.2 MapReduce
- 2.2.1 The Map Tasks
- 2.2.2 Grouping by Key
- 2.2.3 The Reduce Tasks
- 2.2.4 Combiners
- 페이지 수
- 513 페이지
- 업로더
- Uni24h
설명
Trích nội dung tài liệu
Mining of Massive Datasets Jure Leskovec Stanford Univ. Anand Rajaraman Milliway Labs Jeffrey D. Ullman Stanford Univ. Copyright c 2010, 2011, 2012, 2013, 2014 Anand Rajaraman, Jure Leskovec, and Jeffrey D. Ullman ii Preface This book evolved from material developed over several years by Anand Rajaraman and Jeff Ullman for a one-quarter course at Stanford. The course CS345A, titled “Web Mining,” was designed as an advanced graduate course, although it has become accessible and interesting to advanced undergraduates. When Jure Leskovec joined the Stanford faculty, we reorganized the material considerably. He introduced a new course CS224W on network analysis and added material to CS345A, which was renumbered CS246. The three authors also introduced a large-scale data-mining project course, CS341. The book now contains material taught in all three courses. What the Book Is About At the highest level of description, this book is about data mining. However, it focuses on data mining of very large amounts of data, that is, data so large it does not fit in main memory. Because of the emphasis on size, many of our examples are about the Web or data derived from the Web. Further, the book takes an algorithmic point of view: data mining is about applying algorithms to data, rather than using data to “train” a machine-learning engine of some sort. The principal topics covered are: 1. Distributed file systems and map-reduce as a tool for creating parallel algorithms that succeed on very large amounts of data. 2. Similarity search, including the key techniques of minhashing and localitysensitive hashing. 3. Data-stream processing and specialized algorithms for dealing with data that arrives so fast it must be processed immediately or lost. 4. The technology of search engines, including Google’s PageRank, link-spam detection, and the hubs-and-authorities approach. 5. Frequent-itemset mining, including association rules, market-baskets, the A-Priori Algorithm and its impr
자주 묻는 질문
이 문서는 무료인가요?
네. “Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey” 문서는 무료입니다. 로그인 후 '다운로드'를 클릭하여 원본 파일을 받으세요.
이 문서는 몇 페이지로 되어 있나요?
이 문서는 513페이지입니다. 다운로드하기 전에 온라인으로 미리 볼 수 있습니다.
다운로드하기 전에 미리 볼 수 있나요?
네. 이 페이지의 온라인 리더를 통해 문서를 미리 본 후 다운로드 여부를 결정할 수 있습니다.
Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey
미리보기 생성 중...
Trích nội dung tài liệu
Mining of Massive Datasets Jure Leskovec Stanford Univ. Anand Rajaraman Milliway Labs Jeffrey D. Ullman Stanford Univ. Copyright c 2010, 2011, 2012, 2013, 2014 Anand Rajaraman, Jure Leskovec, and Jeffrey D. Ullman ii Preface This book evolved from material developed over several years by Anand Rajaraman and Jeff Ullman for a one-quarter course at Stanford. The course CS345A, titled “Web Mining,” was designed as an advanced graduate course, although it has become accessible and interesting to advanced undergraduates. When Jure Leskovec joined the Stanford faculty, we reorganized the material considerably. He introduced a new course CS224W on network analysis and added material to CS345A, which was renumbered CS246. The three authors also introduced a large-scale data-mining project course, CS341. The book now contains material taught in all three courses. What the Book Is About At the highest level of description, this book is about data mining. However, it focuses on data mining of very large amounts of data, that is, data so large it does not fit in main memory. Because of the emphasis on size, many of our examples are about the Web or data derived from the Web. Further, the book takes an algorithmic point of view: data mining is about applying algorithms to data, rather than using data to “train” a machine-learning engine of some sort. The principal topics covered are: 1. Distributed file systems and map-reduce as a tool for creating parallel algorithms that succeed on very large amounts of data. 2. Similarity search, including the key techniques of minhashing and localitysensitive hashing. 3. Data-stream processing and specialized algorithms for dealing with data that arrives so fast it must be processed immediately or lost. 4. The technology of search engines, including Google’s PageRank, link-spam detection, and the hubs-and-authorities approach. 5. Frequent-itemset mining, including association rules, market-baskets, the A-Priori Algorithm and its impr
- 문서명
- Mining of Massive Datasets - Anand Rajaraman, Jure Leskovec, Jeffrey
- 작성자 (문서 내)
- Anand Rajaraman, Jure Leskovec, Jeffrey D. Ullman
- 내용
- Đây là một cuốn sách về khai thác dữ liệu lớn, tập trung vào các thuật toán xử lý dữ liệu không vừa bộ nhớ chính. Sách bao gồm các chủ đề từ hệ thống phân tán, tìm kiếm tương tự, xử lý luồng, công cụ tìm kiếm, đến phân cụm, đồ thị lớn và học máy cho dữ liệu quy mô lớn.
- 목차
- 2 MapReduce and the New Software Stack
- 2.1 Distributed File Systems
- 2.1.1 Physical Organization of Compute Nodes
- 2.1.2 Large-Scale File-System Organization
- 2.2 MapReduce
- 2.2.1 The Map Tasks
- 2.2.2 Grouping by Key
- 2.2.3 The Reduce Tasks
- 2.2.4 Combiners
- 페이지 수
- 513 페이지
- 업로더
- Uni24h
댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!
Ngân hàng đề thi môn: Hệ thống thông tin quản lý
Đề thi môn Cơ sở dữ liệu (kèm Đáp án) - Đại học Sư phạm kỹ thuật
Đề thi và đáp án môn Hệ thống thông tin kế toán
Đề thi và đáp án môn Cấu trúc dữ liệu giải thuật
Đáp án đề thi môn Mạng máy tính - ĐH Công nghệ thông tin (CNTT)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!