LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)
- 페이지 수
- 23
- 형식
- 크기
- 1.5 MB
- 연도
- 2014
- Trường
- Duke University
- 조회수
- 0
- 댓글
- 0
- Lượt tải
- 0
미리보기 생성 중...
Tài liệu slide bài giảng giới thiệu về Apache Spark, trình bày tổng quan tính toán phân tán, so sánh hệ thống dựa trên đĩa và bộ nhớ, các khái niệm RDD, hành động và biến đổi, cùng kiến trúc chương trình Spark.
- 문서명
- LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)
- 학교 / 강의
- Duke University · Big Data
- 작성자 (문서 내)
- Lance Co Ting Keh
- 내용
- Tài liệu này cung cấp cái nhìn tổng quan về Apache Spark, giải thích lý do sử dụng tính toán phân tán và sự khác biệt giữa các hệ thống lưu trữ. Nó giới thiệu các khái niệm cốt lõi của Spark như RDD và các thành phần khác của hệ sinh thái.
- 목차
- Outline
- I. About me
- II. Distributed CompuIng at a High Level
- III. Disk versus Memory based Systems
- IV. Spark Core
- I. Brief background
- II. Benchmarks and Comparisons
- III. What is an RDD
- IV. RDD AcIons and TransformaIons
- V. Caching and SerializaIon
- VI. Anatomy of a Program
- VII. The Spark Family
- Why Distributed CompuIng?
- Issues Arise in Distributed CompuIng
- Finding majority element in a single machine
- Finding majority element in a distributed dataset
- Disk Based vs Memory Based Frameworks
- The rest of the talk
- Spark Background
- 페이지 수
- 23 페이지
- 업로더
- Uni24h
설명
Trích nội dung tài liệu
Apache Spark 101 Lance Co Ting Keh Senior So5ware Engineer, Machine Learning @ Box 1 Outline I. About me II. Distributed CompuIng at a High Level III. Disk versus Memory based Systems IV. Spark Core I. Brief background II. Benchmarks and Comparisons III. What is an RDD IV. RDD AcIons and TransformaIons V. Caching and SerializaIon VI. Anatomy of a Program VII. The Spark Family 2 Why Distributed CompuIng? Divide and Conquer Problem Single machine cannot complete the computaIon at hand SoluIon Parallelize the job and distribute work among a network of machines 3 Issues Arise in Distributed CompuIng View the world from the eyes of a single worker How do I distribute an algorithm? How do I par++on my dataset? How do I maintain a single consistent view of a shared state? How do I recover from machine failures? How do I allocate cluster resources? ….. 4 Finding majority element in a single machine Think distributed List(20, 18, 20, 18, 20) 5 Finding majority element in a distributed dataset Think distributed List(1, 18, 1, 18, 1) List(2, 18, 2, 18, 2) List(3, 18, 3, 18, 3) List(4, 18, 4, 18, 4) List(5, 18, 5, 18, 5) 6 Finding majority element in a distributed dataset Think distributed List(1, 18, 1, 18, 1) List(2, 18, 2, 18, 2) List(3, 18, 3, 1
자주 묻는 질문
이 문서는 무료인가요?
네. “LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)” 문서는 무료입니다. 로그인 후 '다운로드'를 클릭하여 원본 파일을 받으세요.
이 문서는 몇 페이지로 되어 있나요?
이 문서는 23페이지입니다, Big Data 과정용. 다운로드하기 전에 온라인으로 미리 볼 수 있습니다.
다운로드하기 전에 미리 볼 수 있나요?
네. 이 페이지의 온라인 리더를 통해 문서를 미리 본 후 다운로드 여부를 결정할 수 있습니다.
LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)
미리보기 생성 중...
Trích nội dung tài liệu
Apache Spark 101 Lance Co Ting Keh Senior So5ware Engineer, Machine Learning @ Box 1 Outline I. About me II. Distributed CompuIng at a High Level III. Disk versus Memory based Systems IV. Spark Core I. Brief background II. Benchmarks and Comparisons III. What is an RDD IV. RDD AcIons and TransformaIons V. Caching and SerializaIon VI. Anatomy of a Program VII. The Spark Family 2 Why Distributed CompuIng? Divide and Conquer Problem Single machine cannot complete the computaIon at hand SoluIon Parallelize the job and distribute work among a network of machines 3 Issues Arise in Distributed CompuIng View the world from the eyes of a single worker How do I distribute an algorithm? How do I par++on my dataset? How do I maintain a single consistent view of a shared state? How do I recover from machine failures? How do I allocate cluster resources? ….. 4 Finding majority element in a single machine Think distributed List(20, 18, 20, 18, 20) 5 Finding majority element in a distributed dataset Think distributed List(1, 18, 1, 18, 1) List(2, 18, 2, 18, 2) List(3, 18, 3, 18, 3) List(4, 18, 4, 18, 4) List(5, 18, 5, 18, 5) 6 Finding majority element in a distributed dataset Think distributed List(1, 18, 1, 18, 1) List(2, 18, 2, 18, 2) List(3, 18, 3, 1
- 문서명
- LanceIntroSpark (12) (Tính toán phân tán Apache Spark) (Tiếng Anh)
- 학교 / 강의
- Duke University · Big Data
- 작성자 (문서 내)
- Lance Co Ting Keh
- 내용
- Tài liệu này cung cấp cái nhìn tổng quan về Apache Spark, giải thích lý do sử dụng tính toán phân tán và sự khác biệt giữa các hệ thống lưu trữ. Nó giới thiệu các khái niệm cốt lõi của Spark như RDD và các thành phần khác của hệ sinh thái.
- 목차
- Outline
- I. About me
- II. Distributed CompuIng at a High Level
- III. Disk versus Memory based Systems
- IV. Spark Core
- I. Brief background
- II. Benchmarks and Comparisons
- III. What is an RDD
- IV. RDD AcIons and TransformaIons
- V. Caching and SerializaIon
- VI. Anatomy of a Program
- VII. The Spark Family
- Why Distributed CompuIng?
- Issues Arise in Distributed CompuIng
- Finding majority element in a single machine
- Finding majority element in a distributed dataset
- Disk Based vs Memory Based Frameworks
- The rest of the talk
- Spark Background
- 페이지 수
- 23 페이지
- 업로더
- Uni24h
댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 10)
Advanced Big Data Analytics - Phân tích dữ liệu lớn nâng cao (Lecture 6)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 3)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 4)
Big Data Analytics - Phân tích dữ liệu lớn (Lecture 5)
Tổng hợp Đề Toán 5 - Luyện thi vào Lớp 6 - CLB EMath
Bài giảng vật lý đại cương (Chương 3) - Đỗ Ngọc Uấn
Chương 8.Nguyên tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang

댓글 (0)
댓글이 없습니다. 첫 댓글을 남겨보세요!