Finding Similar Sets (Kỹ thuật tìm tập tương tự)
- 페이지 수
- 33
- 형식
- 크기
- 1.8 MB
- 연도
- 2019
- Trường
- University of Crete
- 조회수
- 0
- 댓글
- 0
- Lượt tải
- 0
미리보기 생성 중...
Slide bài giảng giới thiệu các kỹ thuật tìm tập tương tự bao gồm shingling, minhashing và locality-sensitive hashing.
- 문서명
- Finding Similar Sets (Kỹ thuật tìm tập tương tự)
- 학교 / 강의
- University of Crete · Khai phá dữ liệu
- 작성자 (문서 내)
- Vassilis Christophides
- 내용
- Tài liệu trình bày các phương pháp Shingling, Minhashing và Locality-sensitive Hashing để tìm kiếm các tài liệu tương tự. Nó giải thích cách biểu diễn tài liệu bằng các chuỗi con (shingles), nén chúng bằng hàm băm, và sử dụng các chữ ký để xác định các cặp tài liệu có khả năng giống nhau một cách hiệu quả.
- 목차
- Motivation
- Comparing Documents for Near Duplicates
- Main Issues
- Three Essential Techniques for Detecting Similar Documents
- Shingles
- Shingle Size
- Shingles: Compression Option
- Thought Question
- 페이지 수
- 33 페이지
- 업로더
- Uni24h
설명
Trích nội dung tài liệu
Fall 2019 Finding Similar Sets Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 Fall 2019 Motivation Many Web-mining problems can be expressed as finding “similar” sets: Pages with similar words, e.g., for classification by topic NetFlix users with similar tastes in movies for recommendation systems Dual: movies with similar sets of fans Images of related things The best techniques depend on whether you are looking for items that are very similar or only somewhat similar Special cases are easy, e.g., identical documents, or one document contained character-by-character in another General case, where many small pieces of one document appear out of order in another, is very hard 2 1 1 Fall 2019 Comparing Documents for Near Duplicates Applications: Given a body of documents, find pairs of documents with a lot of text in common, e.g.: Mirror Web sites, or approximate mirrors Application: Don’t want to show both in a search Plagiarism, including large quotations Similar news articles at many news sites Application: Cluster articles by “same story” Simple IR approaches are not suited: Document = set of words appearing in document Document = set of “important” words Why? we need to account for ordering of words! 3 Fall 2019 Main Issues What is the right representation of the document when we check for similarity? E.g., representing a document as a set of characters will not do (why?) When we have billions of documents, keeping the full text in memory is not an option We need to find a shorter representation How do we do pairwise comparisons of billions of documents? If exact match was the issue it would be ok, can we replicate this idea? 4 2 2 Fall 2019 Three Essential Techniques for Detecting Similar Documents Localitysensitive Hashing Document The set of strings of length k that appear in the document Signatures : short integer vectors that represent the sets,
자주 묻는 질문
이 문서는 무료인가요?
네. “Finding Similar Sets (Kỹ thuật tìm tập tương tự)” 문서는 무료입니다. 로그인 후 '다운로드'를 클릭하여 원본 파일을 받으세요.
이 문서는 몇 페이지로 되어 있나요?
이 문서는 33페이지입니다, Khai phá dữ liệu 과정용. 다운로드하기 전에 온라인으로 미리 볼 수 있습니다.
다운로드하기 전에 미리 볼 수 있나요?
네. 이 페이지의 온라인 리더를 통해 문서를 미리 본 후 다운로드 여부를 결정할 수 있습니다.
Finding Similar Sets (Kỹ thuật tìm tập tương tự)
미리보기 생성 중...
Trích nội dung tài liệu
Fall 2019 Finding Similar Sets Vassilis Christophides christop@csd.uoc.gr http://www.csd.uoc.gr/~hy562 University of Crete, Fall 2019 1 Fall 2019 Motivation Many Web-mining problems can be expressed as finding “similar” sets: Pages with similar words, e.g., for classification by topic NetFlix users with similar tastes in movies for recommendation systems Dual: movies with similar sets of fans Images of related things The best techniques depend on whether you are looking for items that are very similar or only somewhat similar Special cases are easy, e.g., identical documents, or one document contained character-by-character in another General case, where many small pieces of one document appear out of order in another, is very hard 2 1 1 Fall 2019 Comparing Documents for Near Duplicates Applications: Given a body of documents, find pairs of documents with a lot of text in common, e.g.: Mirror Web sites, or approximate mirrors Application: Don’t want to show both in a search Plagiarism, including large quotations Similar news articles at many news sites Application: Cluster articles by “same story” Simple IR approaches are not suited: Document = set of words appearing in document Document = set of “important” words Why? we need to account for ordering of words! 3 Fall 2019 Main Issues What is the right representation of the document when we check for similarity? E.g., representing a document as a set of characters will not do (why?) When we have billions of documents, keeping the full text in memory is not an option We need to find a shorter representation How do we do pairwise comparisons of billions of documents? If exact match was the issue it would be ok, can we replicate this idea? 4 2 2 Fall 2019 Three Essential Techniques for Detecting Similar Documents Localitysensitive Hashing Document The set of strings of length k that appear in the document Signatures : short integer vectors that represent the sets,
- 문서명
- Finding Similar Sets (Kỹ thuật tìm tập tương tự)
- 학교 / 강의
- University of Crete · Khai phá dữ liệu
- 작성자 (문서 내)
- Vassilis Christophides
- 내용
- Tài liệu trình bày các phương pháp Shingling, Minhashing và Locality-sensitive Hashing để tìm kiếm các tài liệu tương tự. Nó giải thích cách biểu diễn tài liệu bằng các chuỗi con (shingles), nén chúng bằng hàm băm, và sử dụng các chữ ký để xác định các cặp tài liệu có khả năng giống nhau một cách hiệu quả.
- 목차
- Motivation
- Comparing Documents for Near Duplicates
- Main Issues
- Three Essential Techniques for Detecting Similar Documents
- Shingles
- Shingle Size
- Shingles: Compression Option
- Thought Question
- 페이지 수
- 33 페이지
- 업로더
- Uni24h
댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!
IoT Data Analytics (Phân tích dữ liệu trong Internet vạn vật)
Big Data Processing and Analytics Intro (Xử lý và phân tích dữ liệu lớn)
Introduction to Scalable Data Analytics using Apache Spark (Phân tích dữ liệu quy mô lớn sử dụng Apache Spark)
Frequent Item Sets Association Rules (Tập phổ biến và Luật kết hợp)
Relational Data Processing on MapReduce (Xử lý dữ liệu quan hệ trên MapReduce)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!