LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)
- ページ数
- 23
- 形式
- サイズ
- 1.5 MB
- 年
- 2014
- Trường
- Duke University
- 閲覧数
- 0
- コメント
- 0
- Lượt tải
- 0
プレビューを生成中...
Tài liệu slide bài giảng giới thiệu về Apache Spark, trình bày tổng quan tính toán phân tán, so sánh hệ thống dựa trên đĩa và bộ nhớ, các khái niệm RDD, hành động và biến đổi, cùng kiến trúc chương trình Spark.
- ドキュメント名
- LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)
- 学校 / コース
- Duke University · Big Data
- 著者(ドキュメント内)
- Lance Co Ting Keh
- 内容
- Tài liệu này cung cấp cái nhìn tổng quan về Apache Spark, giải thích lý do sử dụng tính toán phân tán và sự khác biệt giữa các hệ thống lưu trữ. Nó giới thiệu các khái niệm cốt lõi của Spark như RDD và các thành phần khác của hệ sinh thái.
- 目次
- Outline
- I. About me
- II. Distributed CompuIng at a High Level
- III. Disk versus Memory based Systems
- IV. Spark Core
- I. Brief background
- II. Benchmarks and Comparisons
- III. What is an RDD
- IV. RDD AcIons and TransformaIons
- V. Caching and SerializaIon
- VI. Anatomy of a Program
- VII. The Spark Family
- Why Distributed CompuIng?
- Issues Arise in Distributed CompuIng
- Finding majority element in a single machine
- Finding majority element in a distributed dataset
- Disk Based vs Memory Based Frameworks
- The rest of the talk
- Spark Background
- ページ数
- 23 ページ
- アップロード者
- Uni24h
説明
Trích nội dung tài liệu
Apache Spark 101 Lance Co Ting Keh Senior So5ware Engineer, Machine Learning @ Box 1 Outline I. About me II. Distributed CompuIng at a High Level III. Disk versus Memory based Systems IV. Spark Core I. Brief background II. Benchmarks and Comparisons III. What is an RDD IV. RDD AcIons and TransformaIons V. Caching and SerializaIon VI. Anatomy of a Program VII. The Spark Family 2 Why Distributed CompuIng? Divide and Conquer Problem Single machine cannot complete the computaIon at hand SoluIon Parallelize the job and distribute work among a network of machines 3 Issues Arise in Distributed CompuIng View the world from the eyes of a single worker How do I distribute an algorithm? How do I par++on my dataset? How do I maintain a single consistent view of a shared state? How do I recover from machine failures? How do I allocate cluster resources? ….. 4 Finding majority element in a single machine Think distributed List(20, 18, 20, 18, 20) 5 Finding majority element in a distributed dataset Think distributed List(1, 18, 1, 18, 1) List(2, 18, 2, 18, 2) List(3, 18, 3, 18, 3) List(4, 18, 4, 18, 4) List(5, 18, 5, 18, 5) 6 Finding majority element in a distributed dataset Think distributed List(1, 18, 1, 18, 1) List(2, 18, 2, 18, 2) List(3, 18, 3, 1
よくある質問
このドキュメントは無料ですか?
はい。「LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)」は無料です。ログインして「ダウンロード」をクリックするだけで、元のファイルを取得できます。
このドキュメントは何ページありますか?
このドキュメントは 23 ページあります(Big Data コース用)。ダウンロードする前にオンラインでプレビューできます。
ダウンロードする前にプレビューできますか?
はい。このページにあるオンラインリーダーでドキュメントをプレビューし、その後ダウンロードするかどうかを決めることができます。
LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)
プレビューを生成中...
Trích nội dung tài liệu
Apache Spark 101 Lance Co Ting Keh Senior So5ware Engineer, Machine Learning @ Box 1 Outline I. About me II. Distributed CompuIng at a High Level III. Disk versus Memory based Systems IV. Spark Core I. Brief background II. Benchmarks and Comparisons III. What is an RDD IV. RDD AcIons and TransformaIons V. Caching and SerializaIon VI. Anatomy of a Program VII. The Spark Family 2 Why Distributed CompuIng? Divide and Conquer Problem Single machine cannot complete the computaIon at hand SoluIon Parallelize the job and distribute work among a network of machines 3 Issues Arise in Distributed CompuIng View the world from the eyes of a single worker How do I distribute an algorithm? How do I par++on my dataset? How do I maintain a single consistent view of a shared state? How do I recover from machine failures? How do I allocate cluster resources? ….. 4 Finding majority element in a single machine Think distributed List(20, 18, 20, 18, 20) 5 Finding majority element in a distributed dataset Think distributed List(1, 18, 1, 18, 1) List(2, 18, 2, 18, 2) List(3, 18, 3, 18, 3) List(4, 18, 4, 18, 4) List(5, 18, 5, 18, 5) 6 Finding majority element in a distributed dataset Think distributed List(1, 18, 1, 18, 1) List(2, 18, 2, 18, 2) List(3, 18, 3, 1
- ドキュメント名
- LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)
- 学校 / コース
- Duke University · Big Data
- 著者(ドキュメント内)
- Lance Co Ting Keh
- 内容
- Tài liệu này cung cấp cái nhìn tổng quan về Apache Spark, giải thích lý do sử dụng tính toán phân tán và sự khác biệt giữa các hệ thống lưu trữ. Nó giới thiệu các khái niệm cốt lõi của Spark như RDD và các thành phần khác của hệ sinh thái.
- 目次
- Outline
- I. About me
- II. Distributed CompuIng at a High Level
- III. Disk versus Memory based Systems
- IV. Spark Core
- I. Brief background
- II. Benchmarks and Comparisons
- III. What is an RDD
- IV. RDD AcIons and TransformaIons
- V. Caching and SerializaIon
- VI. Anatomy of a Program
- VII. The Spark Family
- Why Distributed CompuIng?
- Issues Arise in Distributed CompuIng
- Finding majority element in a single machine
- Finding majority element in a distributed dataset
- Disk Based vs Memory Based Frameworks
- The rest of the talk
- Spark Background
- ページ数
- 23 ページ
- アップロード者
- Uni24h
コメント (0)
まだコメントはありません。最初のコメントを書きましょう!
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 10)
Advanced Big Data Analytics - Phân tích dữ liệu lớn nâng cao (Lecture 6)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 3)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 4)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 5)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

コメント (0)
まだコメントはありません。最初のコメントを書きましょう!