Distributed Computing with Spark (Tính toán phân tán với Spark) - Reza Zadeh
Tài liệu giới thiệu về tính toán phân tán sử dụng Apache Spark, so sánh mô hình luồng dữ liệu với lập trình mạng truyền thống, hạn chế của MapReduce và kiến trúc RDD của Spark.
Đang tạo bản xem trước...
Distributed Computing with Spark Reza Zadeh Thanks to Matei Zaharia Problem Data growing faster than processing speeds Only solution is to parallelize on large clusters » Wide use in both enterprises and web industry How do we program these things? Outline Data flow vs. traditional network programming Limitations of MapReduce Spark computing engine Machine Learning Example Current State of Spark Ecosystem Built-in Libraries Data flow vs. traditional network programming Traditional Network Programming Message-passing between nodes (e.g. MPI) Very difficult to do at scale: » How to split problem across nodes? Must consider network & data locality » How to deal with failures? (inevitable at scale) » Even worse: stragglers (node not failed, but slow) » Ethernet networking not fast » Have to write programs for each machine Rarely used in commodity datacenters Data Flow Models Restrict the programming interface so that the system can do more automatically Express jobs as graphs of high-level operators » System picks how to split each operator into tasks and where to run each task » Run parts twice fault recovery Map Biggest example: MapReduce Reduce Map Map Reduce Example MapReduce Algorithms Matrix-vector multiplication Power iteration (e.g. PageRank) Gradient descent methods Stochastic SVD Tall skinny QR Many others! Why Use a Data Flow Engine? Ease of programming » High-level functions instead of message passing Wide deployment » More common than MPI, especially “near” data Scalability to very largest clusters » Even HPC world is now concerned about resilience Examples: Pig, Hive, Scalding, Storm Limitations of MapReduce Limitations of MapReduce MapReduce is great at one-pass computation, but inefficient for multi-pass algorithms No efficient primitives for data sharing » State between steps goes to distributed file system » Slow due to replication & disk storage Example: Iterative Apps file system" read file system" write file system" re
… Tải file gốc để đọc toàn bộ tài liệu.
- Tên tài liệu
- Distributed Computing with Spark (Tính toán phân tán với Spark) - Reza Zadeh
- Trường / Môn
- Stanford University · Big Data
- Nội dung
- Tài liệu giới thiệu điện toán phân tán và Spark như một giải pháp cho vấn đề dữ liệu lớn. Nó phân tích hạn chế của MapReduce và trình bày kiến trúc Spark với RDDs, cách xử lý lỗi và các API hỗ trợ.
- Mục lục
- Problem
- Outline
- Data flow vs. traditional network programming
- Traditional Network Programming
- Data Flow Models
- Example MapReduce Algorithms
- Why Use a Data Flow Engine?
- Limitations of MapReduce
- Example: Iterative Apps
- Example: PageRank
- Result
- Spark computing engine
- Spark Computing Engine
- Resilient Distributed Datasets (RDDs)
- Key Idea
- Python, Java, Scala, R
- Fault Tolerance
- Số trang
- 47 trang
- Người đăng
- Uni24h
Câu hỏi thường gặp
Tài liệu này có miễn phí không?
Có. “Distributed Computing with Spark (Tính toán phân tán với Spark) - Reza Zadeh” miễn phí — bạn chỉ cần đăng nhập rồi bấm Tải xuống để lấy file gốc.
Tài liệu dài bao nhiêu trang?
Tài liệu gồm 47 trang, thuộc môn Big Data. Bạn có thể xem trước online trước khi tải.
Tôi có thể xem trước trước khi tải không?
Có. Bạn xem trước tài liệu ngay trên trang này bằng trình đọc online, rồi quyết định tải về.

Bình luận (0)
Chưa có bình luận nào. Hãy là người đầu tiên!