In 2017 a team at Google published a paper titled "Attention Is All You Need," and the architecture it described now powers nearly every major language model. The transformer replaced recurrence with self-attention: every token in a sequence can look at every other token directly, weighing how much each one matters for the current prediction.
That parallel look-up is the key. Older recurrent networks processed words one at a time, which made long-range dependencies hard to learn and training slow. A transformer processes the whole sequence at once, so a pronoun at the end of a paragraph can attend to its antecedent at the beginning without losing the thread.
Core components
- Self-attention layers that compute relationships between all token pairs
- Multi-head attention that runs several attention operations in parallel
- Positional encodings that inject word order information
- Feed-forward blocks and residual connections that stabilize deep stacks
The cost is quadratic: doubling the sequence length quadruples the attention computation. That is why context windows are expensive and why researchers keep proposing efficient variants such as sparse and linear attention.
Comments (3)
Leave a comment