Distributed training of Deep Neural Networks
| dc.contributor.author | Esteves, António | |
| dc.date.accessioned | 2026-09-30T15:38:07Z | |
| dc.date.issued | 2026-07-22 | |
| dc.description.abstract | This report provides a comprehensive and unified study of the principles, methods, and systems that enable scalable training and deployment of large-scale models, with a particular focus on parallelism techniques and Mixture of Experts (MoE) architectures. It combines theoretical insights with practical implementations, offering both a conceptual framework and a hands-on perspective on building efficient distributed systems. The second chapter establishes the foundations of distributed machine learning. The third chapter explores core distributed training strategies, focusing on data parallelism and parameter server architectures. It also presents advanced memory optimization techniques, such as Zero Redundancy Optimizer and Fully Sharded Data Parallel. The fourth chapter focuses on model and computational parallelism, including naive model parallelism, pipeline parallelism, and tensor parallelism. The Mixture of Experts (MoE) paradigm represents a major step forward in scaling neural networks and it is covered on chapter five. Unlike traditional dense models, MoE architectures rely on conditional computation, where only a subset of specialized sub-networks, called experts, is activated for each input. To bridge theory and practice, chapter six includes a comprehensive implementation-focused exploration of MoE models. It revisits the transformer architecture, particularly the decoder-only design used in large language models, and demonstrates how MoE layers can be integrated into this framework. Finally, chapter seven is dedicated to advanced parallelism strategies specifically tailored for MoE architectures, where scalability becomes a multi-dimensional optimization problem. It examines how different parallelism techniques, including data, tensor, pipeline, sequence, and expert parallelism, can be combined into sophisticated hybrid configurations. | eng |
| dc.distribution | n/a | |
| dc.identifier.uri | https://hdl.handle.net/1822/103878 | |
| dc.language.iso | eng | |
| dc.peerreviewed | no | |
| dc.rights | openAccess | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc/4.0/ | |
| dc.subject | Distributed training | |
| dc.subject | Deep neural network | |
| dc.subject | parallelism | |
| dc.subject | Mixture of experts | |
| dc.subject | GPU | |
| dc.subject | PyTorch | |
| dc.subject.fos | Engenharia e Tecnologia::Engenharia Eletrotécnica, Eletrónica e Informática | |
| dc.subject.ods | Indústria, inovação e infraestruturas | |
| dc.title | Distributed training of Deep Neural Networks | eng |
| dc.type | report | |
| dspace.entity.type | Publication |
Ficheiros
Pacote original
1 - 1 de 1
A carregar...
- Nome:
- distributed_training.pdf
- Tamanho:
- 36.43 MB
- Formato:
- Adobe Portable Document Format