EN - FR - DE - ES - IT - PT -

LexiconDream

· Tokenization

Tokenization

Before a language model can read anything, the text has to be chopped into pieces it can count. Those pieces are tokens, and tokenization is the chopping. A token might be a whole word, a fragment like "ing" or "pre", a punctuation mark, or even a single character. The model never sees raw text; it sees a sequence of integer IDs mapped to those tokens.

Different schemes make different trade-offs. Word-level tokenization keeps meaning intact but explodes into millions of rare words and struggles with typos. Character-level tokenization has a tiny vocabulary but produces very long sequences. Subword methods such as byte-pair encoding sit in between: they start with characters and repeatedly merge the most frequent pairs until a target vocabulary size is reached.

Why it matters

A sentence in English might use 20 tokens; the same sentence in Thai or Hindi can take 60 or more under a Western-centric tokenizer. That imbalance affects speed, price, and accuracy for billions of speakers.

Comments

No comments yet. Be the first to share a thought.

Leave a comment