EN - FR - DE - ES - IT - PT -

LexiconDream

🤖 Transformer

A neural network architecture based on self-attention mechanisms.

Transformer

In 2017 a team at Google published a paper titled "Attention Is All You Need," and the architecture it described now powers nearly every major language model. The transformer replaced recurrence with self-attention: every token in a sequence can look at every other token directly, weighing how much each one matters for the current prediction.

That parallel look-up is the key. Older recurrent networks processed words one at a time, which made long-range dependencies hard to learn and training slow. A transformer processes the whole sequence at once, so a pronoun at the end of a paragraph can attend to its antecedent at the beginning without losing the thread.

Core components

The cost is quadratic: doubling the sequence length quadruples the attention computation. That is why context windows are expensive and why researchers keep proposing efficient variants such as sparse and linear attention.

Comments (3)

  1. Lineman Joe
    Transformers are everywhere on poles and pads. Most people never notice them until the power goes out.
  2. Nina B.
    So a transformer changes voltage but not frequency? I want to make sure I understand the basics.
  3. Alex M.
    I always confused electrical transformers with the movie franchise. This page cleared that up.

Leave a comment