NN training research paper (Bài nghiên cứu về huấn luyện mạng nơ-ron NN)
- Pages
- 8
- Format
- Taille
- 1.6 MB
- Année
- 2009
- Vues
- 0
- Commentaires
- 0
- Lượt tải
- 0
Bài báo nghiên cứu về khó khăn trong việc huấn luyện mạng nơ-ron sâu truyền thẳng, phân tích ảnh hưởng của hàm kích hoạt và đề xuất phương pháp khởi tạo mới.
Foire aux questions
Ce document est-il gratuit ?
Oui. « NN training research paper (Bài nghiên cứu về huấn luyện mạng nơ-ron NN) » est gratuit — il suffit de vous connecter et de cliquer sur Télécharger pour obtenir le fichier original.
Combien de pages compte ce document ?
Le document contient 8 pages. Vous pouvez le prévisualiser en ligne avant de le télécharger.
Puis-je prévisualiser avant de télécharger ?
Oui. Vous pouvez prévisualiser ce document directement sur cette page avec le lecteur en ligne, puis décider de le télécharger ou non.
- Nom du document
- NN training research paper (Bài nghiên cứu về huấn luyện mạng nơ-ron NN)
- Auteur (dans le document)
- Xavier Glorot Yoshua Bengio
- Contenu
- Bài báo phân tích lý do mạng nơ-ron sâu khó huấn luyện với khởi tạo ngẫu nhiên, chỉ ra vấn đề của hàm sigmoid và đề xuất phương pháp khởi tạo mới để cải thiện tốc độ hội tụ.
- Table des matières
- 1
- Deep Neural Networks
- Pages
- 8 pages
- Téléversé par
- Uni24h
Génération de l'aperçu...
Description
Understanding the difficulty of training deep feedforward neural networks Xavier Glorot Yoshua Bengio DIRO, Université de Montréal, Montréal, Québec, Canada Abstract Whereas before 2006 it appears that deep multilayer neural networks were not successfully trained, since then several algorithms have been shown to successfully train them, with experimental results showing the superiority of deeper vs less deep architectures. All these experimental results were obtained with new initialization or training mechanisms. Our objective here is to understand better why standard gradient descent from random initialization is doing so poorly with deep neural networks, to better understand these recent relative successes and help design better algorithms in the future. We first observe the influence of the non-linear activations functions. We find that the logistic sigmoid activation is unsuited for deep networks with random initialization because of its mean value, which can drive especially the top hidden layer into saturation. Surprisingly, we find that saturated units can move out of saturation by themselves, albeit slowly, and explaining the plateaus sometimes seen when training neural networks. We find that a new non-linearity that saturates less can often be beneficial. Finally, we study how activations and gradients vary across layers and during training, with the idea that training may be more difficult when the singular values of the Jacobian associated with each layer are far from 1. Based on these considerations, we propose a new initialization scheme that brings substantially faster convergence. 1 Deep Neural Networks Deep learning methods aim at learning feature hierarchies with features from higher levels of the hierarchy formed by the composition of lower level features. They include Appearing in Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS) 2010, Chia Laguna Resort, Sardinia, Italy. Volume 9 of JMLR: W
NN training research paper (Bài nghiên cứu về huấn luyện mạng nơ-ron NN)
Génération de l'aperçu...
Understanding the difficulty of training deep feedforward neural networks Xavier Glorot Yoshua Bengio DIRO, Université de Montréal, Montréal, Québec, Canada Abstract Whereas before 2006 it appears that deep multilayer neural networks were not successfully trained, since then several algorithms have been shown to successfully train them, with experimental results showing the superiority of deeper vs less deep architectures. All these experimental results were obtained with new initialization or training mechanisms. Our objective here is to understand better why standard gradient descent from random initialization is doing so poorly with deep neural networks, to better understand these recent relative successes and help design better algorithms in the future. We first observe the influence of the non-linear activations functions. We find that the logistic sigmoid activation is unsuited for deep networks with random initialization because of its mean value, which can drive especially the top hidden layer into saturation. Surprisingly, we find that saturated units can move out of saturation by themselves, albeit slowly, and explaining the plateaus sometimes seen when training neural networks. We find that a new non-linearity that saturates less can often be beneficial. Finally, we study how activations and gradients vary across layers and during training, with the idea that training may be more difficult when the singular values of the Jacobian associated with each layer are far from 1. Based on these considerations, we propose a new initialization scheme that brings substantially faster convergence. 1 Deep Neural Networks Deep learning methods aim at learning feature hierarchies with features from higher levels of the hierarchy formed by the composition of lower level features. They include Appearing in Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS) 2010, Chia Laguna Resort, Sardinia, Italy. Volume 9 of JMLR: W
Lire le document entier
- Nom du document
- NN training research paper (Bài nghiên cứu về huấn luyện mạng nơ-ron NN)
- Auteur (dans le document)
- Xavier Glorot Yoshua Bengio
- Contenu
- Bài báo phân tích lý do mạng nơ-ron sâu khó huấn luyện với khởi tạo ngẫu nhiên, chỉ ra vấn đề của hàm sigmoid và đề xuất phương pháp khởi tạo mới để cải thiện tốc độ hội tụ.
- Table des matières
- 1
- Deep Neural Networks
- Pages
- 8 pages
- Téléversé par
- Uni24h
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !
Ngân hàng đề thi môn: Hệ thống thông tin quản lý
Đề thi môn Cơ sở dữ liệu (kèm Đáp án) - Đại học Sư phạm kỹ thuật
Đề thi và đáp án môn Hệ thống thông tin kế toán
Đề thi và đáp án môn Cấu trúc dữ liệu giải thuật
Đáp án đề thi môn Mạng máy tính - ĐH Công nghệ thông tin (CNTT)
Chương 7.Cơ học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 6.Quang học lượng tử - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 5.Thuyết tương đối - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 4. Tán xạ ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Chương 3.Phân cực ánh sáng - Vật lý đại cương 3 - TS.Nguyễn Thị Trang
Commentaires (0)
Aucun commentaire pour le moment. Soyez le premier !