AIWiki
Malaysia
Back to all articles
Modelsrwkvlinear-attentionsequence-models

RWKV

4 min readUpdated October 2026
RWKV
Type
Hybrid RNN / linear-attention architecture
Developer
Bo Peng (BlinkDL) and the RWKV community
First released
2023 (RWKV-4 paper, EMNLP 2023)
Latest version
RWKV-7 Goose (2025)
Licence
Apache 2.0 (models and code)
Related
Mamba, Transformers, RetNet
RWKV (Receptance Weighted Key Value, pronounced "RwaKuv") is a neural network architecture for sequence modelling that combines the training efficiency of Transformers with the constant-memory inference of recurrent neural networks. It replaces softmax attention with a linear recurrence governed by learned time decay, so generation cost and memory stay flat regardless of context length. RWKV models are released as open weights under the Apache 2.0 licence.

History

RWKV was developed by Bo Peng, who works under the handle BlinkDL, and was first described in the 2023 paper RWKV: Reinventing RNNs for the Transformer Era, published at EMNLP 2023 Findings. The paper showed that RWKV-4 matched similarly sized Transformers on standard language-modelling benchmarks while retaining RNN-style inference, and the architecture was scaled to 14 billion parameters — at the time the largest dense RNN trained.[1]

In 2024 the team published Eagle (RWKV-5) and Finch (RWKV-6), which extend the architecture with multi-headed matrix-valued states and a dynamic recurrence whose decay depends on the input, alongside a 1.12-trillion-token multilingual training corpus and a fast multilingual tokenizer.[2] All models were released under Apache 2.0. A seventh generation, RWKV-7 Goose, arrived in 2025 with a state-update rule inspired by delta-rule memory models, further closing the gap with softmax attention on recall-heavy tasks.[3]

Key Concepts and Technology

RWKV's core is the WKV operator, which computes an attention-like weighted average as a recurrence: past key-value pairs are accumulated with an exponentially decaying weight determined by a learned per-channel time decay, while the current token receives a separate bonus weight. A receptance gate then reads out the accumulated state.[1] Because each step only updates a fixed-size state rather than revisiting the whole sequence, inference runs in O(1) memory and time per token with no KV cache — in contrast to Transformers, whose KV cache grows linearly with context length.

The architecture also uses token shift, a learned interpolation between the current and previous token embeddings that injects local bigram context, and splits each residual block into time-mixing and channel-mixing sub-layers analogous to attention and feed-forward layers. Crucially, the recurrence can be computed as a parallel scan during training, so RWKV keeps the time-parallel training of a Transformer while generating like an RNN.[2]

The trade-off is retrieval fidelity: full softmax attention remains stronger at precise recall of arbitrary tokens deep in a long context (the needle-in-a-haystack setting), while RWKV excels at throughput and memory efficiency. This has driven a broader industry shift toward hybrid architectures that combine attention with recurrent or state-space components.

Applications or Impact

RWKV's constant-memory profile makes it attractive for on-device inference, long-running agents and high-throughput serving where KV-cache growth dominates cost. Its Apache 2.0 terms permit commercial use without the licence conditions attached to some other open model families, which has made it a reference architecture in the open-weight ecosystem alongside Llama and Qwen. RWKV also demonstrated that viable Transformer alternatives can be trained and maintained by a community-driven project rather than only by large labs.

>See Also

🇲🇾Malaysian Context

🇲🇾 Relevance to Malaysia: RWKV's flat memory footprint suits deployments where GPU budgets are tight — Malaysian SMEs, universities and public-sector projects running models on modest hardware or at the edge can serve long conversations without paying the KV-cache memory tax. That aligns with the National AI Office's emphasis on practical, affordable AI adoption and with edge-AI use cases in Penang's electronics and embedded-systems manufacturing base.

Because RWKV weights are Apache 2.0, local developers can fine-tune and ship them commercially with clear licence provenance — a consideration for organisations that must document software licences for audit. Its multilingual training corpus and small model sizes also make it a candidate for on-premises assistants in sectors where PDPA requires personal data to stay on Malaysian infrastructure.

References

  1. ↑Peng, B. et al. (2023). RWKV: Reinventing RNNs for the Transformer Era. EMNLP 2023 Findings. https://arxiv.org/abs/2305.13048
  2. ↑Peng, B. et al. (2024). Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence. https://arxiv.org/abs/2404.05892
  3. ↑RWKV Community. The RWKV Architecture Wiki. https://wiki.rwkv.ac/
  4. ↑GitHub. BlinkDL/RWKV-LM — official RWKV implementation. https://github.com/BlinkDL/RWKV-LM