Apache Spark Performance Troubleshooting at Scale Challenges, Tools and Methods (Khắc phục hiệu năng của Apache Spark ở quy mô lớn) (Tiếng Anh)
- Pages
- 48
- Format
- PPTX
- Taille
- 3 MB
- Année
- 2000
- Vues
- 0
- Commentaires
- 0
- Lượt tải
- 0
Bài trình bày về các thách thức, công cụ và phương pháp để khắc phục hiệu năng của Apache Spark ở quy mô lớn.
Foire aux questions
Ce document est-il gratuit ?
Oui. « Apache Spark Performance Troubleshooting at Scale Challenges, Tools and Methods (Khắc phục hiệu năng của Apache Spark ở quy mô lớn) (Tiếng Anh) » est gratuit — il suffit de vous connecter et de cliquer sur Télécharger pour obtenir le fichier original.
Combien de pages compte ce document ?
Le document contient 48 pages. Vous pouvez le prévisualiser en ligne avant de le télécharger.
Puis-je prévisualiser avant de télécharger ?
Oui. Vous pouvez prévisualiser ce document directement sur cette page avec le lecteur en ligne, puis décider de le télécharger ou non.
- Nom du document
- Apache Spark Performance Troubleshooting at Scale Challenges, Tools and Methods (Khắc phục hiệu năng của Apache Spark ở quy mô lớn) (Tiếng Anh)
- Auteur (dans le document)
- Luca Canali
- Contenu
- Tài liệu thảo luận về những thách thức khi khắc phục sự cố hiệu năng Apache Spark ở quy mô lớn, giới thiệu các công cụ và phương pháp đo lường, phân tích để đưa ra các giải pháp khả thi. Bài thuyết trình nhấn mạnh tầm quan trọng của việc hiểu rõ hệ thống và sử dụng công cụ phù hợp để đạt được hiệu quả tối ưu.
- Table des matières
- Ce document n'a pas de table des matières claire.
- Pages
- 48 pages
- Téléversé par
- Uni24h
Génération de l'aperçu...
Description
Apache Spark Performance Troubleshooting at Scale: Challenges, Tools and Methods Luca Canali, CERN #EUdev2 About Luca Computing engineer and team lead at CERN IT Hadoop and Spark service, database services Joined CERN in 2005 17+ years of experience with database services Performance, architecture, tools, internals Sharing information: blog, notes, code @LucaCanaliDB – http://cern.ch/canali #EUdev2 2 CERN and the Large Hadron Collider Largest and most powerful particle accelerator #EUdev2 3 Apache Spark @ Spark is a popular component for data processing Deployed on four production Hadoop/YARN clusters Aggregated capacity (2017): ~1500 physical cores, 11 PB Adoption is growing. Key projects involving Spark: Analytics for accelerator controls and logging Monitoring use cases, this includes use of Spark streaming Analytics on aggregated logs Explorations on the use of Spark for high energy physics Link: http://cern.ch/canali/docs/BigData_Solutions_at_CERN_KT_Forum_20170929.pdf #EUdev2 4 Motivations for This Work Understanding Spark workloads Understanding technology (where are the bottlenecks, how much do Spark jobs scale, etc?) Capacity planning: benchmark platforms Provide our users with a range of monitoring tools Measurements and troubleshooting Spark SQL Structured data in Parquet for data analytics Spark-ROOT (project on using Spark for physics data) #EUdev2 5 Outlook of This Talk Topic is vast, I will just share some ideas and lessons learned How to approach performance troubleshooting, benchmarking and relevant methods Data sources and tools to measure Spark workloads, challenges at scale Examples and lessons learned with some key tools #EUdev2 6 Challenges Just measuring performance metrics is easy Producing actionable insights requires effort and preparation Methods on how to approach troubleshooting performance How to gather relevant data Need to use the right tools, poss
Apache Spark Performance Troubleshooting at Scale Challenges, Tools and Methods (Khắc phục hiệu năng của Apache Spark ở quy mô lớn) (Tiếng Anh)
Génération de l'aperçu...
Apache Spark Performance Troubleshooting at Scale: Challenges, Tools and Methods Luca Canali, CERN #EUdev2 About Luca Computing engineer and team lead at CERN IT Hadoop and Spark service, database services Joined CERN in 2005 17+ years of experience with database services Performance, architecture, tools, internals Sharing information: blog, notes, code @LucaCanaliDB – http://cern.ch/canali #EUdev2 2 CERN and the Large Hadron Collider Largest and most powerful particle accelerator #EUdev2 3 Apache Spark @ Spark is a popular component for data processing Deployed on four production Hadoop/YARN clusters Aggregated capacity (2017): ~1500 physical cores, 11 PB Adoption is growing. Key projects involving Spark: Analytics for accelerator controls and logging Monitoring use cases, this includes use of Spark streaming Analytics on aggregated logs Explorations on the use of Spark for high energy physics Link: http://cern.ch/canali/docs/BigData_Solutions_at_CERN_KT_Forum_20170929.pdf #EUdev2 4 Motivations for This Work Understanding Spark workloads Understanding technology (where are the bottlenecks, how much do Spark jobs scale, etc?) Capacity planning: benchmark platforms Provide our users with a range of monitoring tools Measurements and troubleshooting Spark SQL Structured data in Parquet for data analytics Spark-ROOT (project on using Spark for physics data) #EUdev2 5 Outlook of This Talk Topic is vast, I will just share some ideas and lessons learned How to approach performance troubleshooting, benchmarking and relevant methods Data sources and tools to measure Spark workloads, challenges at scale Examples and lessons learned with some key tools #EUdev2 6 Challenges Just measuring performance metrics is easy Producing actionable insights requires effort and preparation Methods on how to approach troubleshooting performance How to gather relevant data Need to use the right tools, poss
Lire le document entier
- Nom du document
- Apache Spark Performance Troubleshooting at Scale Challenges, Tools and Methods (Khắc phục hiệu năng của Apache Spark ở quy mô lớn) (Tiếng Anh)
- Auteur (dans le document)
- Luca Canali
- Contenu
- Tài liệu thảo luận về những thách thức khi khắc phục sự cố hiệu năng Apache Spark ở quy mô lớn, giới thiệu các công cụ và phương pháp đo lường, phân tích để đưa ra các giải pháp khả thi. Bài thuyết trình nhấn mạnh tầm quan trọng của việc hiểu rõ hệ thống và sử dụng công cụ phù hợp để đạt được hiệu quả tối ưu.
- Table des matières
- Ce document n'a pas de table des matières claire.
- Pages
- 48 pages
- Téléversé par
- Uni24h
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !
Ngân hàng đề thi môn: Hệ thống thông tin quản lý
Đề thi môn Cơ sở dữ liệu (kèm Đáp án) - Đại học Sư phạm kỹ thuật
Đề thi và đáp án môn Hệ thống thông tin kế toán
Đề thi và đáp án môn Cấu trúc dữ liệu giải thuật
Đáp án đề thi môn Mạng máy tính - ĐH Công nghệ thông tin (CNTT)
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !