Distributed Computing with Spark (Tính toán phân tán với Spark) - Reza Zadeh
Tài liệu giới thiệu về tính toán phân tán sử dụng Apache Spark, so sánh mô hình luồng dữ liệu với lập trình mạng truyền thống, hạn chế của MapReduce và kiến trúc RDD của Spark.
미리보기 생성 중...
Distributed Computing with Spark Reza Zadeh Thanks to Matei Zaharia Problem Data growing faster than processing speeds Only solution is to parallelize on large clusters » Wide use in both enterprises and web industry How do we program these things? Outline Data flow vs. traditional network programming Limitations of MapReduce Spark computing engine Machine Learning Example Current State of Spark Ecosystem Built-in Libraries Data flow vs. traditional network programming Traditional Network Programming Message-passing between nodes (e.g. MPI) Very difficult to do at scale: » How to split problem across nodes? Must consider network & data locality » How to deal with failures? (inevitable at scale) » Even worse: stragglers (node not failed, but slow) » Ethernet networking not fast » Have to write programs for each machine Rarely used in commodity datacenters Data Flow Models Restrict the programming interface so that the system can do more automatically Express jobs as graphs of high-level operators » System picks how to split each operator into tasks and where to run each task » Run parts twice fault recovery Map Biggest example: MapReduce Reduce Map Map Reduce Example MapReduce Algorithms Matrix-vector multiplication Power iteration (e.g. PageRank) Gradient descent methods Stochastic SVD Tall skinny QR Many others! Why Use a Data Flow Engine? Ease of programming » High-level functions instead of message passing Wide deployment » More common than MPI, especially “near” data Scalability to very largest clusters » Even HPC world is now concerned about resilience Examples: Pig, Hive, Scalding, Storm Limitations of MapReduce Limitations of MapReduce MapReduce is great at one-pass computation, but inefficient for multi-pass algorithms No efficient primitives for data sharing » State between steps goes to distributed file system » Slow due to replication & disk storage Example: Iterative Apps file system" read file system" write file system" re
… 전체 문서를 읽으려면 원본 파일을 다운로드하세요.
- 문서명
- Distributed Computing with Spark (Tính toán phân tán với Spark) - Reza Zadeh
- 학교 / 강의
- Stanford University · Big Data
- 내용
- Tài liệu giới thiệu điện toán phân tán và Spark như một giải pháp cho vấn đề dữ liệu lớn. Nó phân tích hạn chế của MapReduce và trình bày kiến trúc Spark với RDDs, cách xử lý lỗi và các API hỗ trợ.
- 목차
- Problem
- Outline
- Data flow vs. traditional network programming
- Traditional Network Programming
- Data Flow Models
- Example MapReduce Algorithms
- Why Use a Data Flow Engine?
- Limitations of MapReduce
- Example: Iterative Apps
- Example: PageRank
- Result
- Spark computing engine
- Spark Computing Engine
- Resilient Distributed Datasets (RDDs)
- Key Idea
- Python, Java, Scala, R
- Fault Tolerance
- 페이지 수
- 47 페이지
- 업로더
- Uni24h
자주 묻는 질문
이 문서는 무료인가요?
네. “Distributed Computing with Spark (Tính toán phân tán với Spark) - Reza Zadeh” 문서는 무료입니다. 로그인 후 '다운로드'를 클릭하여 원본 파일을 받으세요.
이 문서는 몇 페이지로 되어 있나요?
이 문서는 47페이지입니다, Big Data 과정용. 다운로드하기 전에 온라인으로 미리 볼 수 있습니다.
다운로드하기 전에 미리 볼 수 있나요?
네. 이 페이지의 온라인 리더를 통해 문서를 미리 본 후 다운로드 여부를 결정할 수 있습니다.

댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!