keyword
Transformer
A Transformer is a deep learning model architecture designed to process sequential data, such as text, by relying on a self-attention mechanism to weigh the relationships between all elements in a sequence simultaneously. Unlike earlier recurrent architectures that processed inputs step by step, the Transformer analyzes entire sequences in parallel, significantly improving computational efficiency and scalability during training. The core architecture consists of stacked layers featuring multi-head self-attention, position-wise feed-forward networks, residual connections, and layer normalization, which together enable the model to capture complex contextual dependencies across long distances. While initially developed for natural language processing tasks such as machine translation, the Transformer has become the foundational framework for modern large language models, computer vision systems, and multimodal artificial intelligence.
1 item

