Spark (07) (Cách làm việc với Spark và thư viện học máy)
- Seiten
- 49
- Định dạng
- Dung lượng
- 1 MB
- Trường
- University of Hildesheim
- Aufrufe
- 0
- Kommentare
- 0
- Lượt tải
- 0
Vorschau wird generiert...
Slide bài giảng số 07 về Apache Spark, trình bày về Resilient Distributed Datasets, cách làm việc với Spark và thư viện học máy MLLib.
- Dokumentenname
- Spark (07) (Cách làm việc với Spark và thư viện học máy)
- Schule / Kurs
- University of Hildesheim · Big Data
- Autor (im Dokument)
- Lars Schmidt-Thieme
- Inhalt
- Tài liệu giới thiệu về Apache Spark và RDDs trong bối cảnh tính toán phân tán cho Big Data Analytics. Nó giải thích cách Spark xử lý dữ liệu phân tán và đảm bảo tính chịu lỗi thông qua việc lưu trữ lịch sử biến đổi dữ liệu.
- Inhaltsverzeichnis
- 0. Introduction
- A. Parallel Computing
- A.1 Threads
- A.2 Message Passing Interface (MPI)
- A.3 Graphical Processing Units (GPUs)
- B. Distributed Storage
- B.1 Distributed File Systems
- B.2 Partioning of Relational Databases
- B.3 NoSQL Databases
- C. Distributed Computing Environments
- C.1 Map-Reduce
- C.2 Resilient Distributed Datasets (Spark)
- C.3 Computational Graphs (TensorFlow)
- D. Distributed Machine Learning Algorithms
- D.1 Distributed Stochastic Gradient Descent
- D.2 Distributed Matrix Factorization
- Questions and Answers
- 1. Introduction
- 2. Apache Spark
- 3. Working with Spark
- 4. MLLib: Machine Learning with Spark
- Seiten
- 49 Seiten
- Hochgeladen von
- Uni24h
Beschreibung
Trích nội dung tài liệu
Big Data Analytics Big Data Analytics C. Distributed Computing Environments / C.2. Resilient Distributed Datasets: Apache Spark Lars Schmidt-Thieme Information Systems and Machine Learning Lab (ISMLL) Institute for Computer Science University of Hildesheim, Germany Lars Schmidt-Thieme, Information Systems and Machine Learning Lab (ISMLL), University of Hildesheim, Germany 1 / 35 Big Data Analytics Syllabus Tue. 9.4. (1) 0. Introduction Tue. 16.4. Tue. 23.4. Tue. 30.4. (2) (3) (4) A. Parallel Computing A.1 Threads A.2 Message Passing Interface (MPI) A.3 Graphical Processing Units (GPUs) Tue. 7.5. Tue. 14.5. Tue. 21.5. (5) (6) (7) B. Distributed Storage B.1 Distributed File Systems B.2 Partioning of Relational Databases B.3 NoSQL Databases Tue. 28.5. Tue. 4.6. Tue. 11.6. Tue. 18.6. (8) (9) (10) C. Distributed Computing Environments C.1 Map-Reduce C.2 Resilient Distributed Datasets (Spark) Pentecoste Break — C.3 Computational Graphs (TensorFlow) Tue. 25.6. Tue. 2.7. (11) (12) D. Distributed Machine Learning Algorithms D.1 Distributed Stochastic Gradient Descent D.2 Distributed Matrix Factorization Tue. 9.7. (13) Questions and Answers Lars Schmidt-Thieme, Information Systems and Machine Learning Lab (ISMLL), University of Hildesheim, Germany 1 / 35 Big Data Analytics Outline 1. Introduction 2. Apache Spark 3. Working with Spark 4. MLLib: Machine Learning with Spark Lars Schmidt-Thieme, Information Systems and Machine Learning Lab (ISMLL), University of Hildesheim, Germany 1 / 35 Big Data Analytics 1. Introduction Outline 1. Introduction 2. Apache Spark 3. Working with Spark 4. MLLib: Machine Learning with Spark Lars Schmidt-Thieme, Information Systems and Machine Learning Lab (ISMLL), University of Hildesheim, Germany 1 / 35 Big Data Analytics 1. Introduction Core Idea To implement fault-tolerance for primary/original data: I replication: I partition large data into parts I store each part several times on different servers
Häufig gestellte Fragen
Ist dieses Dokument kostenlos?
Ja. „Spark (07) (Cách làm việc với Spark và thư viện học máy)“ ist kostenlos — melden Sie sich einfach an und klicken Sie auf Herunterladen, um die Originaldatei zu erhalten.
Wie viele Seiten hat dieses Dokument?
Das Dokument hat 49 Seiten, für den Kurs Big Data. Sie können es vor dem Herunterladen online in der Vorschau ansehen.
Kann ich vor dem Herunterladen eine Vorschau ansehen?
Ja. Sie können sich dieses Dokument direkt auf dieser Seite im Online-Reader ansehen und dann entscheiden, ob Sie es herunterladen möchten.
Spark (07) (Cách làm việc với Spark và thư viện học máy)
Vorschau wird generiert...
Trích nội dung tài liệu
Big Data Analytics Big Data Analytics C. Distributed Computing Environments / C.2. Resilient Distributed Datasets: Apache Spark Lars Schmidt-Thieme Information Systems and Machine Learning Lab (ISMLL) Institute for Computer Science University of Hildesheim, Germany Lars Schmidt-Thieme, Information Systems and Machine Learning Lab (ISMLL), University of Hildesheim, Germany 1 / 35 Big Data Analytics Syllabus Tue. 9.4. (1) 0. Introduction Tue. 16.4. Tue. 23.4. Tue. 30.4. (2) (3) (4) A. Parallel Computing A.1 Threads A.2 Message Passing Interface (MPI) A.3 Graphical Processing Units (GPUs) Tue. 7.5. Tue. 14.5. Tue. 21.5. (5) (6) (7) B. Distributed Storage B.1 Distributed File Systems B.2 Partioning of Relational Databases B.3 NoSQL Databases Tue. 28.5. Tue. 4.6. Tue. 11.6. Tue. 18.6. (8) (9) (10) C. Distributed Computing Environments C.1 Map-Reduce C.2 Resilient Distributed Datasets (Spark) Pentecoste Break — C.3 Computational Graphs (TensorFlow) Tue. 25.6. Tue. 2.7. (11) (12) D. Distributed Machine Learning Algorithms D.1 Distributed Stochastic Gradient Descent D.2 Distributed Matrix Factorization Tue. 9.7. (13) Questions and Answers Lars Schmidt-Thieme, Information Systems and Machine Learning Lab (ISMLL), University of Hildesheim, Germany 1 / 35 Big Data Analytics Outline 1. Introduction 2. Apache Spark 3. Working with Spark 4. MLLib: Machine Learning with Spark Lars Schmidt-Thieme, Information Systems and Machine Learning Lab (ISMLL), University of Hildesheim, Germany 1 / 35 Big Data Analytics 1. Introduction Outline 1. Introduction 2. Apache Spark 3. Working with Spark 4. MLLib: Machine Learning with Spark Lars Schmidt-Thieme, Information Systems and Machine Learning Lab (ISMLL), University of Hildesheim, Germany 1 / 35 Big Data Analytics 1. Introduction Core Idea To implement fault-tolerance for primary/original data: I replication: I partition large data into parts I store each part several times on different servers
- Dokumentenname
- Spark (07) (Cách làm việc với Spark và thư viện học máy)
- Schule / Kurs
- University of Hildesheim · Big Data
- Autor (im Dokument)
- Lars Schmidt-Thieme
- Inhalt
- Tài liệu giới thiệu về Apache Spark và RDDs trong bối cảnh tính toán phân tán cho Big Data Analytics. Nó giải thích cách Spark xử lý dữ liệu phân tán và đảm bảo tính chịu lỗi thông qua việc lưu trữ lịch sử biến đổi dữ liệu.
- Inhaltsverzeichnis
- 0. Introduction
- A. Parallel Computing
- A.1 Threads
- A.2 Message Passing Interface (MPI)
- A.3 Graphical Processing Units (GPUs)
- B. Distributed Storage
- B.1 Distributed File Systems
- B.2 Partioning of Relational Databases
- B.3 NoSQL Databases
- C. Distributed Computing Environments
- C.1 Map-Reduce
- C.2 Resilient Distributed Datasets (Spark)
- C.3 Computational Graphs (TensorFlow)
- D. Distributed Machine Learning Algorithms
- D.1 Distributed Stochastic Gradient Descent
- D.2 Distributed Matrix Factorization
- Questions and Answers
- 1. Introduction
- 2. Apache Spark
- 3. Working with Spark
- 4. MLLib: Machine Learning with Spark
- Seiten
- 49 Seiten
- Hochgeladen von
- Uni24h
Kommentare (0)
Noch keine Kommentare. Seien Sie der Erste!
Stream (11) (Xử lý luồng dữ liệu) - Julian M. Kunkel
Krone (09) (Sự phát triển của dữ liệu) (Tiếng Anh)
Parallel mf (09) (Thuật toán phân tán phân tích ma trận dữ liệu lớn)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 5)
NoSQL db (06) (Cơ sở dữ liệu NoSQL)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

Kommentare (0)
Noch keine Kommentare. Seien Sie der Erste!