ML Map Reduce (Lecture 11) (Máy học với MapReduce)
Génération de l'aperçu...
Slide bài giảng về Machine Learning với MapReduce, bao gồm các phần về K-Means Clustering và EM-Algorithm.
Description
Machine Learning with MapReduce K-Means Clustering 3 How to MapReduce K-Means? Given K, assign the first K random points to be the initial cluster centers Assign subsequent points to the closest cluster using the supplied distance measure Compute the centroid of each cluster and iterate the previous step until the cluster centers converge within delta Run a final pass over the points to cluster them for output K-Means Map/Reduce Design Driver Runs multiple iteration jobs using mapper+combiner+reducer Runs final clustering job using only mapper Mapper Configure: Single file containing encoded Clusters Input: File split containing encoded Vectors Output: Vectors keyed by nearest cluster Combiner Input: Vectors keyed by nearest cluster Output: Cluster centroid vectors keyed by “cluster” Reducer (singleton) Input: Cluster centroid vectors Output: Single file containing Vectors keyed by cluster Mapper - mapper has k centers in memory. Input Key-value pair (each input data point x). Find the index of the closest of the k centers (call it iClosest). Emit: (key,value) = (iClosest, x) Reducer(s) – Input (key,value) Key = index of center Value = iterator over input data points closest to ith center At each key value, run through the iterator and average all the Corresponding input data points. Emit: (index of center, new center) Improved Version: Calculate partial sums in mappers Mapper - mapper has k centers in memory. Running through one input data point at a time (call it x). Find the index of the closest of the k centers (call it iClosest). Accumulate sum of inputs segregated into K groups depending on which center is closest. Emit: ( , partial sum) Or Emit(index, partial sum) Reducer – accumulate partial sums and Emit with index or without EM-Algorithm What is MLE? Given A sample X={X1, …, Xn} A vector of parameters θ We define Likelihood of the data: P(X | θ) Log-likelihood of the data: L(θ)=log P(X|θ)
Résumé IA
- Nom du document
- ML Map Reduce (Lecture 11) (Máy học với MapReduce)
- École / Cours
- University of Hamburg · Big Data
- Contenu
- Tài liệu giới thiệu cách triển khai thuật toán K-Means Clustering và EM-Algorithm bằng MapReduce. Nó giải thích các bước thiết kế Map, Combiner, Reducer cho K-Means và trình bày nền tảng lý thuyết của EM-Algorithm, bao gồm MLE và cách xử lý dữ liệu ẩn.
- Table des matières
- K-Means Clustering
- How to MapReduce K-Means?
- K-Means Map/Reduce Design
- Mapper - mapper has k centers in memory.
- Improved Version: Calculate partial sums in mappers
- EM-Algorithm
- What is MLE?
- MLE (cont)
- An easy case
- An easy case (cont)
- Basic setting in EM
- The basic EM strategy
- The log-likelihood function
- The iterative approach for MLE
- Pages
- 51 pages
- Téléversé par
- Uni24h
Foire aux questions
Ce document est-il gratuit ?
Oui. « ML Map Reduce (Lecture 11) (Máy học với MapReduce) » est gratuit — il suffit de vous connecter et de cliquer sur Télécharger pour obtenir le fichier original.
Combien de pages compte ce document ?
Le document contient 51 pages, pour le cours Big Data. Vous pouvez le prévisualiser en ligne avant de le télécharger.
Puis-je prévisualiser avant de télécharger ?
Oui. Vous pouvez prévisualiser ce document directement sur cette page avec le lecteur en ligne, puis décider de le télécharger ou non.
ML Map Reduce (Lecture 11) (Máy học với MapReduce)
Génération de l'aperçu...
Machine Learning with MapReduce K-Means Clustering 3 How to MapReduce K-Means? Given K, assign the first K random points to be the initial cluster centers Assign subsequent points to the closest cluster using the supplied distance measure Compute the centroid of each cluster and iterate the previous step until the cluster centers converge within delta Run a final pass over the points to cluster them for output K-Means Map/Reduce Design Driver Runs multiple iteration jobs using mapper+combiner+reducer Runs final clustering job using only mapper Mapper Configure: Single file containing encoded Clusters Input: File split containing encoded Vectors Output: Vectors keyed by nearest cluster Combiner Input: Vectors keyed by nearest cluster Output: Cluster centroid vectors keyed by “cluster” Reducer (singleton) Input: Cluster centroid vectors Output: Single file containing Vectors keyed by cluster Mapper - mapper has k centers in memory. Input Key-value pair (each input data point x). Find the index of the closest of the k centers (call it iClosest). Emit: (key,value) = (iClosest, x) Reducer(s) – Input (key,value) Key = index of center Value = iterator over input data points closest to ith center At each key value, run through the iterator and average all the Corresponding input data points. Emit: (index of center, new center) Improved Version: Calculate partial sums in mappers Mapper - mapper has k centers in memory. Running through one input data point at a time (call it x). Find the index of the closest of the k centers (call it iClosest). Accumulate sum of inputs segregated into K groups depending on which center is closest. Emit: ( , partial sum) Or Emit(index, partial sum) Reducer – accumulate partial sums and Emit with index or without EM-Algorithm What is MLE? Given A sample X={X1, …, Xn} A vector of parameters θ We define Likelihood of the data: P(X | θ) Log-likelihood of the data: L(θ)=log P(X|θ)
Lire le document entier
- Nom du document
- ML Map Reduce (Lecture 11) (Máy học với MapReduce)
- École / Cours
- University of Hamburg · Big Data
- Contenu
- Tài liệu giới thiệu cách triển khai thuật toán K-Means Clustering và EM-Algorithm bằng MapReduce. Nó giải thích các bước thiết kế Map, Combiner, Reducer cho K-Means và trình bày nền tảng lý thuyết của EM-Algorithm, bao gồm MLE và cách xử lý dữ liệu ẩn.
- Table des matières
- K-Means Clustering
- How to MapReduce K-Means?
- K-Means Map/Reduce Design
- Mapper - mapper has k centers in memory.
- Improved Version: Calculate partial sums in mappers
- EM-Algorithm
- What is MLE?
- MLE (cont)
- An easy case
- An easy case (cont)
- Basic setting in EM
- The basic EM strategy
- The log-likelihood function
- The iterative approach for MLE
- Pages
- 51 pages
- Téléversé par
- Uni24h
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !