AIWiki
Malaysia
Back to all articles
AI Foundationstokenizationbyte pair encodingnlp

Tokenization (Natural Language Processing)

4 min readUpdated September 2026
Tokenization
Type
Natural language processing technique
Purpose
Convert text into units a model can process numerically
Common methods
Byte-pair encoding, WordPiece, SentencePiece
Key metric
Fertility — the average number of tokens per word
Emerged
Statistical NLP era; standard for language models since 2018
Related
Large language models, embeddings, context windows
Tokenization is the process of breaking text into smaller units, called tokens, that a language model can convert into numbers and process. Every modern AI language system — from chatbots to translation tools — reads and writes through a tokenizer, and the design of that tokenizer quietly determines how much text fits in a model's context window, how much an AI feature costs to run, and how well the system handles languages other than English.[1][2]

Background

Early natural language processing systems split text on spaces and punctuation, treating words as the atomic unit. That approach breaks down for languages without spaces between words, for compound words, and for the many forms a single root word can take. The neural era replaced fixed rules with subword tokenization: algorithms such as byte-pair encoding, originally developed as a data-compression technique, learn from a training corpus which character sequences occur together often enough to be merged into single tokens. WordPiece, used in BERT, and SentencePiece, designed to be language-agnostic, followed; byte-level schemes adopted by GPT-family models guarantee that any text — including emoji, mixed scripts and code — can be represented as a sequence of tokens.[1][3]

How It Works

A tokenizer is trained separately from the model itself. It builds a vocabulary — typically tens of thousands to a few hundred thousand entries — of the most useful words, word fragments and bytes, then applies it consistently at both training and run time. Frequent English words such as "the" or "model" usually occupy a single token; an unusual name might be split into three or four fragments.

The efficiency of that split is measured by fertility, the average number of tokens a word or sentence requires. English text commonly averages roughly one token per short word, but the same content in other languages can consume several times as many tokens. An often-cited illustration from OpenAI's developer community notes that for some tokenizers a message in one language can require ten to twenty times more tokens than a comparable English message; studies of Indian and Southeast Asian languages have found similar gaps, and tokenizer choice has become a recognised axis of multilingual AI research.[2][3][4]

Applications and Impact

Because commercial AI services bill by the token, tokenization has direct economic consequences: text that tokenizes heavily costs more to process, fills context windows faster and can make a chatbot in a lower-resource language several times more expensive to operate than the English equivalent. Developers building retrieval-augmented systems, translators or local-language assistants therefore treat tokenizer behaviour as an engineering variable — choosing models with more efficient coverage, compressing prompts, or fine-tuning on domain text so that common phrases become compact tokens.[1][4]

Tokenization also shapes fairness. A model whose tokenizer fragments a language unnaturally sees that language in rougher pieces, which can degrade quality on tasks such as translation, summarisation and sentiment analysis. Researchers have proposed frameworks for evaluating and adapting tokenizers for under-served language families, including a 2026 conceptual framework examining efficiency trade-offs in adapting large language models for ASEAN languages.[1][2]

>See Also

🇲🇾Malaysian Context

🇲🇾 Tokenization is central to Malaysia's efforts to make AI work well in Bahasa Melayu. Studies of Southeast Asian language adaptation note that tokenizers trained predominantly on English allocate comparatively few vocabulary entries to Malay words and affixes, so Malay prompts typically consume more tokens — and more cost — than English prompts of the same meaning, a burden that grows with code-switching into Manglish or regional dialects.[1][3]

Malaysian model builders have responded by treating language coverage as a first-class goal. ILMU, developed by YTL AI Labs, was built to understand Bahasa Melayu, Manglish and regional dialects, while Mesolitica's MaLLaM family was pretrained from scratch on a Malay-centric corpus. For local developers, the practical implications are familiar: prefer models with demonstrated Malay token efficiency when cost matters, test tokenizer behaviour before committing to an API provider, and budget more generously for Malay-language context windows.[1]

References

  1. ↑Cureus. (2026). Beyond the Tokenization Bottleneck: A Conceptual Framework for Efficiency-Plasticity Trade-offs in ASEAN Large Language Model Adaptation. https://www.cureusjournals.com/articles/13958-beyond-the-tokenization-bottleneck-a-conceptual-framework-for-efficiency-plasticity-trade-offs-in-asean-large-language-model-adaptation
  2. ↑Tamang, S. and Bora, D. J. (2024). Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages. arXiv:2411.12240. https://arxiv.org/html/2411.12240v1
  3. ↑OpenAI Developer Community. All languages are NOT created (tokenized) equal. https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407
  4. ↑TokenCalculator.ai. Most Token-Efficient Languages for LLMs, Ranked and Priced. https://tokencalculator.ai/most-token-efficient-languages-for-llms-ranked-priced/