transformer

A neural network architecture based on the attention mechanism that underlies large language models.

3 articles
Last mentioned

The transformer is a neural network architecture built on the attention mechanism, which computes how strongly each part of an input relates to the others. Google researchers proposed it in the 2017 paper Attention Is All You Need. Because it can process sequential data such as text in parallel, it became the foundation of large language models such as GPT, Claude and Gemini, as well as many image and speech models.

This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.

Articles covering this entry

He adds that today's transformers cannot inspect their own intermediate reasoning once it has passed through their layers, so they struggle to know what they do…

Collison goes further and calls computer use the next major step in AI, following transformers, large language models and reinforcement learning.

He calls the "cardinal sin" of transformers the mixing of world knowledge and reasoning procedures in the same weights.


© 2026 AIPOST. All rights reserved.

AIPOST is an AI publication covering practical AI, AI security, performance, startups, health, ethics and industry news. No account is needed, and our privacy policy explains how we handle personal information.