LoRA and QLoRA
Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning method that freezes a pretrained model's original weights and learns small low-rank update matrices in selected layers. QLoRA combines LoRA with a frozen, quantized base model and backpropagates gradients through that representation into the adapters. Hugging Face PEFT exposes LoRA through a maintained configuration and adapter API. QLoRA is therefore a specific memory-saving training recipe built on LoRA, not a synonym for every quantized model or adapter method.
Origin and context
Hu and colleagues introduced LoRA in 2021 as a way to adapt large models without storing or updating a full set of task-specific parameters. Their experiments inserted trainable rank-decomposition matrices while keeping pretrained weights fixed. In 2023, Dettmers and colleagues presented QLoRA, which trained LoRA adapters through a frozen 4-bit quantized language model and added NormalFloat 4, double quantization, and paged optimizers. Hugging Face subsequently incorporated configurable LoRA-family support into PEFT, including QLoRA-style targeting of all linear layers.
Why it matters
These methods reduce the trainable-state and memory burden of adapting a large model. The LoRA paper reported far fewer trainable parameters and lower GPU memory use than full fine-tuning in its evaluated settings, while QLoRA reported fine-tuning a 65-billion-parameter model on one 48 GB GPU. A supported library implementation makes the methods usable through repeatable adapter configurations. The savings concern adaptation and adapter weights; they do not erase the cost of obtaining, loading, evaluating, or serving the base model.
Example
A team adapting one language model for two tasks can keep one frozen base checkpoint and train a separate LoRA adapter for each task. In Hugging Face PEFT, a LoraConfig selects the rank, scaling, dropout, and target modules; QLoRA-style training can target all linear layers while using a quantized base representation. An implementation may later load adapters dynamically or merge compatible LoRA weights. Quantizing a model only for serving, without training low-rank adapters, is not QLoRA.
Maturity and evidence
LoRA and QLoRA merit maturity 4 as established techniques. Independent research groups published detailed methods and experiments, QLoRA explicitly builds on LoRA, and Hugging Face PEFT documents a maintained implementation with configurable targeting, adapter loading, merging, and multiple LoRA variants. That is concrete adoption evidence beyond the originating papers. The rating describes concept and implementation maturity, not uniform performance across every model; a maturity 5 rating would require stronger cross-stack predictability and long-term compatibility evidence.
Limits and open questions
Parameter efficiency does not guarantee that an adapted model matches full fine-tuning on every task. Results depend on target layers, rank, data quality, optimization, quantization choices, and the base model. QLoRA's memory and quality findings are experimental results from specified model families and hardware, not universal guarantees. Library support also evolves, so teams should pin compatible versions and test merging and quantization behavior. A combined entry should preserve the methods' distinct definitions and avoid treating all PEFT approaches as LoRA.
Related terms
References
- LoRA: Low-Rank Adaptation of Large Language ModelsMicrosoft Research / arXiv · 2021-06-17 · class A
- QLoRA: Efficient Finetuning of Quantized LLMsUniversity of Washington / arXiv · 2023-05-23 · class A
- PEFT LoRA package referenceHugging Face · 2026 · class A
Last updated: 2026-08-27
This term is also covered in the Skills Atlas as lora qlora skill.
This term is also covered in the Skills Atlas as hugging face peft skill.
This term is also covered in the Skills Atlas as llm fine tuning skill.