- Type
- Hybrid RNN / linear-attention architecture
- Developer
- Bo Peng (BlinkDL) and the RWKV community
- First released
- 2023 (RWKV-4 paper, EMNLP 2023)
- Latest version
- RWKV-7 Goose (2025)
- Licence
- Apache 2.0 (models and code)
- Related
- Mamba, Transformers, RetNet
- Type
- Hybrid RNN / linear-attention architecture
- Developer
- Bo Peng (BlinkDL) and the RWKV community
- First released
- 2023 (RWKV-4 paper, EMNLP 2023)
- Latest version
- RWKV-7 Goose (2025)
- Licence
- Apache 2.0 (models and code)
- Related
- Mamba, Transformers, RetNet
History
RWKV was developed by Bo Peng, who works under the handle BlinkDL, and was first described in the 2023 paper RWKV: Reinventing RNNs for the Transformer Era, published at EMNLP 2023 Findings. The paper showed that RWKV-4 matched similarly sized Transformers on standard language-modelling benchmarks while retaining RNN-style inference, and the architecture was scaled to 14 billion parameters — at the time the largest dense RNN trained.[1]
In 2024 the team published Eagle (RWKV-5) and Finch (RWKV-6), which extend the architecture with multi-headed matrix-valued states and a dynamic recurrence whose decay depends on the input, alongside a 1.12-trillion-token multilingual training corpus and a fast multilingual tokenizer.[2] All models were released under Apache 2.0. A seventh generation, RWKV-7 Goose, arrived in 2025 with a state-update rule inspired by delta-rule memory models, further closing the gap with softmax attention on recall-heavy tasks.[3]
Key Concepts and Technology
RWKV's core is the WKV operator, which computes an attention-like weighted average as a recurrence: past key-value pairs are accumulated with an exponentially decaying weight determined by a learned per-channel time decay, while the current token receives a separate bonus weight. A receptance gate then reads out the accumulated state.[1] Because each step only updates a fixed-size state rather than revisiting the whole sequence, inference runs in O(1) memory and time per token with no KV cache — in contrast to Transformers, whose KV cache grows linearly with context length.
The architecture also uses token shift, a learned interpolation between the current and previous token embeddings that injects local bigram context, and splits each residual block into time-mixing and channel-mixing sub-layers analogous to attention and feed-forward layers. Crucially, the recurrence can be computed as a parallel scan during training, so RWKV keeps the time-parallel training of a Transformer while generating like an RNN.[2]
The trade-off is retrieval fidelity: full softmax attention remains stronger at precise recall of arbitrary tokens deep in a long context (the needle-in-a-haystack setting), while RWKV excels at throughput and memory efficiency. This has driven a broader industry shift toward hybrid architectures that combine attention with recurrent or state-space components.
Applications or Impact
RWKV's constant-memory profile makes it attractive for on-device inference, long-running agents and high-throughput serving where KV-cache growth dominates cost. Its Apache 2.0 terms permit commercial use without the licence conditions attached to some other open model families, which has made it a reference architecture in the open-weight ecosystem alongside Llama and Qwen. RWKV also demonstrated that viable Transformer alternatives can be trained and maintained by a community-driven project rather than only by large labs.
>See Also
🇲🇾 Relevance to Malaysia: RWKV's flat memory footprint suits deployments where GPU budgets are tight — Malaysian SMEs, universities and public-sector projects running models on modest hardware or at the edge can serve long conversations without paying the KV-cache memory tax. That aligns with the National AI Office's emphasis on practical, affordable AI adoption and with edge-AI use cases in Penang's electronics and embedded-systems manufacturing base.
Because RWKV weights are Apache 2.0, local developers can fine-tune and ship them commercially with clear licence provenance — a consideration for organisations that must document software licences for audit. Its multilingual training corpus and small model sizes also make it a candidate for on-premises assistants in sectors where PDPA requires personal data to stay on Malaysian infrastructure.
References
- ↑Peng, B. et al. (2023). RWKV: Reinventing RNNs for the Transformer Era. EMNLP 2023 Findings. https://arxiv.org/abs/2305.13048
- ↑Peng, B. et al. (2024). Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence. https://arxiv.org/abs/2404.05892
- ↑RWKV Community. The RWKV Architecture Wiki. https://wiki.rwkv.ac/
- ↑GitHub. BlinkDL/RWKV-LM — official RWKV implementation. https://github.com/BlinkDL/RWKV-LM