Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
正在生成预览...
Slide bài giảng về tính toán trong bộ nhớ với Spark, bao gồm kiến trúc, mô hình dữ liệu RDD, biến chia sẻ và các ví dụ.
描述
In-Memory Computation with Spark Lecture BigData Analytics Julian M. Kunkel julian.kunkel@googlemail.com University of Hamburg / German Climate Computing Center (DKRZ) 2017-01-20 Disclaimer: Big Data software is constantly updated, code samples may be outdated. Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary Outline 1 Concepts 2 Architecture 3 Computation 4 Managing Jobs 5 Examples 6 Higher-Level Abstractions 7 Summary Julian M. Kunkel Lecture BigData Analytics, 2016 2 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary In-Memory Computation/Processing/Analytics [26] In-memory processing: Processing data stored in memory (database) Advantage: No slow I/O necessary ⇒ fast response times Disadvantages Data must fit in the memory of the distributed storage/database Additional persistency (with asynchronous flushing) usually required Fault-tolerance is mandatory BI-Solution: SAP Hana Big data approaches: Apache Spark, Apache Flink Julian M. Kunkel Lecture BigData Analytics, 2016 3 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary Overview to Spark [10, 12] In-memory processing (and storage) engine Load data from HDFS, Cassandra, HBase Resource management via. YARN, Mesos, Spark, Amazon EC2 ⇒ It can use Hadoop but also works standalone! Task scheduling and monitoring Rich APIs APIs for Java, Scala, Python, R Thrift JDBC/ODBC server for SQL High-level domain-specific tools/languages Advanced APIs simplify typical computation tasks Interactive shells with tight integration spark-shell: Scala (object-oriented functional language running on JVM) pyspark: Python sparkR: R (basic support) Execution in either local (single node) or cluster mode Julian M. Kunkel Lecture BigData Analytics, 2016 4 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions
AI 摘要
- 文档名称
- Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
- 学校 / 课程
- University of Hamburg · Big Data
- 内容
- Bài giảng giới thiệu Apache Spark như một công cụ xử lý dữ liệu lớn trong bộ nhớ, tập trung vào các khái niệm như RDDs, kiến trúc, tính toán và quản lý tác vụ. Tài liệu cũng đề cập đến các API và cách thức thực thi của Spark.
- 目录
- Concepts
- Architecture
- Computation
- Managing Jobs
- Examples
- Higher-Level Abstractions
- Summary
- 页数
- 54 页
- 上传者
- Uni24h
常见问题
此文档免费吗?
是的。“Tính toán trong bộ nhớ với Spark - Julian M. Kunkel”是免费的 — 只需登录并点击“下载”即可获取原始文件。
这份文档有多少页?
该文档共有 54 页,适用于课程 Big Data。您可以在下载前进行在线预览。
我可以在下载前预览吗?
是的。您可以通过在线阅读器直接在本页面预览此文档,然后再决定是否下载。
Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
正在生成预览...
In-Memory Computation with Spark Lecture BigData Analytics Julian M. Kunkel julian.kunkel@googlemail.com University of Hamburg / German Climate Computing Center (DKRZ) 2017-01-20 Disclaimer: Big Data software is constantly updated, code samples may be outdated. Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary Outline 1 Concepts 2 Architecture 3 Computation 4 Managing Jobs 5 Examples 6 Higher-Level Abstractions 7 Summary Julian M. Kunkel Lecture BigData Analytics, 2016 2 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary In-Memory Computation/Processing/Analytics [26] In-memory processing: Processing data stored in memory (database) Advantage: No slow I/O necessary ⇒ fast response times Disadvantages Data must fit in the memory of the distributed storage/database Additional persistency (with asynchronous flushing) usually required Fault-tolerance is mandatory BI-Solution: SAP Hana Big data approaches: Apache Spark, Apache Flink Julian M. Kunkel Lecture BigData Analytics, 2016 3 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions Summary Overview to Spark [10, 12] In-memory processing (and storage) engine Load data from HDFS, Cassandra, HBase Resource management via. YARN, Mesos, Spark, Amazon EC2 ⇒ It can use Hadoop but also works standalone! Task scheduling and monitoring Rich APIs APIs for Java, Scala, Python, R Thrift JDBC/ODBC server for SQL High-level domain-specific tools/languages Advanced APIs simplify typical computation tasks Interactive shells with tight integration spark-shell: Scala (object-oriented functional language running on JVM) pyspark: Python sparkR: R (basic support) Execution in either local (single node) or cluster mode Julian M. Kunkel Lecture BigData Analytics, 2016 4 / 53 Concepts Architecture Computation Managing Jobs Examples Higher-Level Abstractions
阅读全文
- 文档名称
- Tính toán trong bộ nhớ với Spark - Julian M. Kunkel
- 学校 / 课程
- University of Hamburg · Big Data
- 内容
- Bài giảng giới thiệu Apache Spark như một công cụ xử lý dữ liệu lớn trong bộ nhớ, tập trung vào các khái niệm như RDDs, kiến trúc, tính toán và quản lý tác vụ. Tài liệu cũng đề cập đến các API và cách thức thực thi của Spark.
- 目录
- Concepts
- Architecture
- Computation
- Managing Jobs
- Examples
- Higher-Level Abstractions
- Summary
- 页数
- 54 页
- 上传者
- Uni24h
评论 (0)
暂无评论。快来抢沙发吧!
Neumann (mối quan hệ giữa Exascale Computing và Big Data) - Philipp Neumann
Intro to Mapreduce (02) (Giới thiệu về MapReduce và Hadoop) (Tiếng Anh)
GPUs (04) (Xử lý song song và bộ xử lý đồ họa)
Parallel computing (01) (Tính toán song song)
Neo4j (08) (Xử lý đồ thị với Neo4j) - BigData Analytics - Julian M. Kunkel
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
评论 (0)
暂无评论。快来抢沙发吧!