AIWiki
Malaysia
Back to all articles
AI Foundationsqlorafine-tuningquantization

QLoRA

4 min readUpdated September 2026
QLoRA
Type
Parameter-efficient fine-tuning technique
Developed by
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer (University of Washington)
Introduced
May 2023 (arXiv preprint; NeurIPS 2023)
Key ideas
4-bit NormalFloat, double quantization, paged optimizers
Notable result
Fine-tuning a 65B-parameter model on a single 48GB GPU
Related
LoRA, quantization, fine-tuning
QLoRA (Quantized Low-Rank Adaptation) is a memory-efficient fine-tuning method that allows very large language models to be adapted to new tasks on modest hardware. Introduced in a May 2023 paper by Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer at the University of Washington, QLoRA quantises a pretrained model to 4-bit precision, freezes it, and trains small low-rank adapter weights on top — a combination that reduced the memory needed to fine-tune a 65-billion-parameter model to a single 48GB GPU while matching the task performance of full 16-bit fine-tuning.[1][2]

Background

Fine-tuning large language models traditionally required enough accelerator memory to hold the model weights, gradients and optimiser state in 16-bit precision — for a 65B-parameter model, well over 780GB of memory, beyond all but the largest clusters. Two research lines changed that equation in the early 2020s. LoRA (Low-Rank Adaptation, 2021) showed that most of the benefit of fine-tuning could be captured by training small rank-decomposition matrices added to the frozen original weights, cutting trainable parameters by orders of magnitude. Quantisation research meanwhile showed that model weights could be stored at 8-bit and 4-bit precision with limited quality loss for inference. QLoRA combined the two and showed that fine-tuning through a 4-bit quantised base model could recover full 16-bit fine-tuning quality.[1][3]

Key Concepts and Technology

QLoRA introduces three main techniques. 4-bit NormalFloat (NF4) is a quantisation data type that is information-theoretically optimal for the zero-centred normal distribution typical of pretrained weights, yielding better accuracy than plain 4-bit integers or floats. Double quantization compresses the quantization constants themselves, saving roughly 0.37 bits per parameter and allowing larger models to fit into fixed GPU memory budgets. Paged optimizers use NVIDIA unified memory to page optimiser state to CPU memory during memory spikes, preventing out-of-memory failures during gradient checkpointing. During training, QLoRA keeps one storage data type (usually NF4) and one computation data type (16-bit BrainFloat); gradients backpropagate through the frozen 4-bit model into the adapters only, which are stored and trained in higher precision. The paper's results showed 4-bit QLoRA matching 16-bit full fine-tuning and 16-bit LoRA on standard academic benchmarks, and enabled 33B models to be fine-tuned on a single 24GB GPU. The method was integrated into the Hugging Face transformers and bitsandbytes libraries within weeks of publication, and the accompanying Guanaco model family was fine-tuned on a single professional GPU.[1][2][4]

Applications and Impact

QLoRA became one of the most widely used fine-tuning methods of the open-model era. Researchers, startups and hobbyists use it to adapt open-weight models on consumer graphics cards such as the NVIDIA RTX 3090 and 4090, which changed the economics of customisation: instead of renting GPU clusters by the hour, a team can fine-tune a domain- or language-specific model on a single workstation. Its influence is visible across the fine-tuning ecosystem, including subsequent tools and frameworks that adopted QLoRA-style pipelines for instruction tuning, domain adaptation and low-resource language work. The method is also studied in academic settings as a subject in its own right, with follow-up research on quantisation formats, adapter placement and memory management.[2][4]

>See Also

🇲🇾Malaysian Context

QLoRA matters to Malaysian teams because it collapses the cost of adapting large models from a cluster-scale budget to a workstation-scale one. Universities, research groups and startups in Malaysia have used QLoRA-style pipelines to fine-tune open-weight models for Bahasa Melayu and mixed-language tasks, for domain use cases in Islamic finance, legal and healthcare settings, and for experiments in low-resource language preservation — work that is otherwise difficult to fund when GPU hours are priced in foreign currency. The technique also supports data-sovereignty strategies: models can be fine-tuned on premises, keeping sensitive institutional data within the organisation in line with obligations under the Personal Data Protection Act 2010, rather than uploading it to third-party APIs. Local training providers and university AI programmes increasingly include parameter-efficient fine-tuning in their certifications, reflecting demand for engineers who can customise models rather than only call them.[1]

References

  1. ↑Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. arXiv:2305.14314. https://arxiv.org/abs/2305.14314
  2. ↑Hugging Face. (2023). Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA. https://huggingface.co/blog/4bit-transformers-bitsandbytes
  3. ↑Hu, E., et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685. https://arxiv.org/abs/2106.09685
  4. ↑Dettmers, T., et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. Advances in Neural Information Processing Systems (NeurIPS 2023). https://proceedings.neurips.cc/paper_files/paper/2023/hash/1feb87871436031bdc0f2beaa62a049b-Abstract-Conference.html