transformer
A neural network architecture based on the attention mechanism that underlies large language models.
The transformer is a neural network architecture built on the attention mechanism, which computes how strongly each part of an input relates to the others. Google researchers proposed it in the 2017 paper Attention Is All You Need. Because it can process sequential data such as text in parallel, it became the foundation of large language models such as GPT, Claude and Gemini, as well as many image and speech models.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Collison goes further and calls computer use the next major step in AI, following transformers, large language models and reinforcement learning.
He calls the "cardinal sin" of transformers the mixing of world knowledge and reasoning procedures in the same weights.