Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
- Seiten
- 54
- Định dạng
- Dung lượng
- 826 KB
- Năm
- 2016
- Trường
- University of Hamburg
- Aufrufe
- 0
- Kommentare
- 0
- Lượt tải
- 0
Vorschau wird generiert...
Slide bài giảng về tính toán trong bộ nhớ với Spark, bao gồm kiến trúc, mô hình dữ liệu RDD, biến chia sẻ và các ví dụ.
- Dokumentenname
- Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
- Schule / Kurs
- University of Hamburg · Big Data
- Inhalt
- Bài giảng giới thiệu Apache Spark như một công cụ xử lý dữ liệu lớn trong bộ nhớ, tập trung vào các khái niệm như RDDs, kiến trúc, tính toán và quản lý tác vụ. Tài liệu cũng đề cập đến các API và cách thức thực thi của Spark.
- Inhaltsverzeichnis
- Concepts
- Architecture
- Computation
- Managing Jobs
- Examples
- Higher-Level Abstractions
- Summary
- Seiten
- 54 Seiten
- Hochgeladen von
- Uni24h
Beschreibung
Trích nội dung tài liệu
In-Memory Computation with Spark Lecture BigData Analytics Julian M. Kunkel julian.kunkel@googlemail.com University of Hamburg / German Climate Computing Center (DKRZ) 2017-01-20 Disclaimer: Big Data software is constantly updated, code samples may be outdated. Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary Outline 1 Concepts 2 Architecture 3 Computation 4 Managing Jobs 5 Examples 6 Higher-Level Abstractions 7 Summary Julian M. Kunkel Lecture BigData Analytics, 2016 2 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary In-Memory Computation/Processing/Analytics [26] In-memory processing: Processing data stored in memory (database) Advantage: No slow I/O necessary ⇒ fast response times Disadvantages Data must fit in the memory of the distributed storage/database Additional persistency (with asynchronous flushing) usually required Fault-tolerance is mandatory BI-Solution: SAP Hana Big data approaches: Apache Spark, Apache Flink Julian M. Kunkel Lecture BigData Analytics, 2016 3 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary Overview to Spark [10, 12] In-memory processing (and storage) engine Load data from HDFS, Cassandra, HBase Resource management via. YARN, Mesos, Spark, Amazon EC2 ⇒ It can use Hadoop but also works standalone! Task scheduling and monitoring Rich APIs APIs for Java, Scala, Python, R Thrift JDBC/ODBC server for SQL High-level domain-specific tools/languages Advanced APIs simplify typical computation tasks Interactive shells with tight integration spark-shell: Scala (object-oriented functional language running on JVM) pyspark: Python sparkR: R (basic support) Execution in either local (single node) or cluster mode Julian M. Kunkel Lecture BigData Analytics, 2016 4 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions
Häufig gestellte Fragen
Ist dieses Dokument kostenlos?
Ja. „Tính toán trong bộ nhớ với Spark - Julian M. Kunkel“ ist kostenlos — melden Sie sich einfach an und klicken Sie auf Herunterladen, um die Originaldatei zu erhalten.
Wie viele Seiten hat dieses Dokument?
Das Dokument hat 54 Seiten, für den Kurs Big Data. Sie können es vor dem Herunterladen online in der Vorschau ansehen.
Kann ich vor dem Herunterladen eine Vorschau ansehen?
Ja. Sie können sich dieses Dokument direkt auf dieser Seite im Online-Reader ansehen und dann entscheiden, ob Sie es herunterladen möchten.
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
Vorschau wird generiert...
Trích nội dung tài liệu
In-Memory Computation with Spark Lecture BigData Analytics Julian M. Kunkel julian.kunkel@googlemail.com University of Hamburg / German Climate Computing Center (DKRZ) 2017-01-20 Disclaimer: Big Data software is constantly updated, code samples may be outdated. Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary Outline 1 Concepts 2 Architecture 3 Computation 4 Managing Jobs 5 Examples 6 Higher-Level Abstractions 7 Summary Julian M. Kunkel Lecture BigData Analytics, 2016 2 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary In-Memory Computation/Processing/Analytics [26] In-memory processing: Processing data stored in memory (database) Advantage: No slow I/O necessary ⇒ fast response times Disadvantages Data must fit in the memory of the distributed storage/database Additional persistency (with asynchronous flushing) usually required Fault-tolerance is mandatory BI-Solution: SAP Hana Big data approaches: Apache Spark, Apache Flink Julian M. Kunkel Lecture BigData Analytics, 2016 3 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary Overview to Spark [10, 12] In-memory processing (and storage) engine Load data from HDFS, Cassandra, HBase Resource management via. YARN, Mesos, Spark, Amazon EC2 ⇒ It can use Hadoop but also works standalone! Task scheduling and monitoring Rich APIs APIs for Java, Scala, Python, R Thrift JDBC/ODBC server for SQL High-level domain-specific tools/languages Advanced APIs simplify typical computation tasks Interactive shells with tight integration spark-shell: Scala (object-oriented functional language running on JVM) pyspark: Python sparkR: R (basic support) Execution in either local (single node) or cluster mode Julian M. Kunkel Lecture BigData Analytics, 2016 4 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions
- Dokumentenname
- Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
- Schule / Kurs
- University of Hamburg · Big Data
- Inhalt
- Bài giảng giới thiệu Apache Spark như một công cụ xử lý dữ liệu lớn trong bộ nhớ, tập trung vào các khái niệm như RDDs, kiến trúc, tính toán và quản lý tác vụ. Tài liệu cũng đề cập đến các API và cách thức thực thi của Spark.
- Inhaltsverzeichnis
- Concepts
- Architecture
- Computation
- Managing Jobs
- Examples
- Higher-Level Abstractions
- Summary
- Seiten
- 54 Seiten
- Hochgeladen von
- Uni24h
Kommentare (0)
Noch keine Kommentare. Seien Sie der Erste!
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 10)
Advanced Big Data Analytics - Phân tích dữ liệu lớn nâng cao (Lecture 6)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 3)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 4)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 5)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

Kommentare (0)
Noch keine Kommentare. Seien Sie der Erste!