☆ 4.5 Article

Layer-Parallel Training of Deep Residual Neural Networks

SIAM JOURNAL ON MATHEMATICS OF DATA SCIENCE (2020)

期刊

SIAM JOURNAL ON MATHEMATICS OF DATA SCIENCE

卷 2, 期 1, 页码 1-23

出版社

SIAM PUBLICATIONS

DOI: 10.1137/19M1247620

关键词

deep learning; residual networks; supervised learning; optimal control; layer-parallelization; parallel-in-time; simultaneous optimization

类别

Mathematics, Applied

资金

U.S. National Science Foundation [DMS 1522599, DMS 1751636]
Sandia National Laboratories
U.S. Department of Energy's National Nuclear Security Administration [DE-NA0003525]

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

Residual neural networks (ResNets) are a promising class of deep neural networks that have shown excellent performance for a number of learning tasks, e.g., image classification and recognition. Mathematically, ResNet architectures can be interpreted as forward Euler discretizations of a nonlinear initial value problem whose time-dependent control variables represent the weights of the neural network. Hence, training a ResNet can be cast as an optimal control problem of the associated dynamical system. For similar time-dependent optimal control problems arising in engineering applications, parallel-in-time methods have shown notable improvements in scalability. This paper demonstrates the use of those techniques for efficient and effective training of ResNets. The proposed algorithms replace the classical (sequential) forward and backward propagation through the network layers with a parallel nonlinear multigrid iteration applied to the layer domain. This adds a new dimension of parallelism across layers that is attractive when training very deep networks. From this basic idea, we derive multiple layer-parallel methods. The most efficient version employs a simultaneous optimization approach where updates to the network parameters are based on inexact gradient information in order to speed up the training process. Using numerical examples from supervised classification, we demonstrate that the new approach achieves a training performance similar to that of traditional methods, but enables layer-parallelism and thus provides speedup over layer-serial methods through greater concurrency.

Layer-Parallel Training of Deep Residual Neural Networks

期刊

SIAM JOURNAL ON MATHEMATICS OF DATA SCIENCE

出版社

SIAM PUBLICATIONS

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Layer-Parallel Training of Deep Residual Neural Networks

期刊

SIAM JOURNAL ON MATHEMATICS OF DATA SCIENCE

出版社

SIAM PUBLICATIONS

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文