Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PEFT is a family of methods that fine-tune a pretrained model by training a relatively small set of added parameters rather than updating every model weight. LoRA is one PEFT method; QLoRA uses LoRA adapters while keeping the base model quantized. That distinction can lower memory pressure, but it does not guarantee that a particular model and training setup will fit a particular GPU.

How PEFT, LoRA, and QLoRA fit together

Full fine-tuning updates the pretrained model’s weights. PEFT instead trains a comparatively small number of added parameters on top of those weights, which can make adaptation less demanding in memory and storage.

LoRA: train adapters, not the full base model

Low-Rank Adaptation (LoRA) adds trainable low-rank adapter parameters to selected parts of a model. During fine-tuning, those adapters are trained while the base weights are not updated. LoRA is therefore a type of PEFT, not a synonym for the whole PEFT family. See the Hugging Face PEFT quantization guide for its documented implementation context.

QLoRA: LoRA over a quantized base

Quantized Low-Rank Adaptation (QLoRA) combines LoRA adapters with a quantized base model. The adapters remain trainable; quantization reduces the precision used to represent the base weights. QLoRA is not a separate alternative to LoRA so much as a way of applying LoRA to a quantized model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the three approaches compare

Approach Weights trained Base model quantized? Memory implications Configuration considerations
Full fine-tuning Pretrained model weights No, not by definition Can require substantial memory because model weights are updated during training. Training setup depends on the model and workload.
LoRA Added low-rank adapter parameters No, not by definition Trains fewer parameters than full fine-tuning; memory use still depends on model and training configuration. Adapter configuration and target modules depend on the model architecture and task.
QLoRA LoRA adapter parameters Yes Quantizing the base can reduce memory pressure; actual requirements vary with the model and configuration. Requires a compatible quantization setup as well as model-appropriate adapter settings.

The available sources do not establish a universal speed, quality, or cost ranking among these methods. A smaller trainable parameter count or quantized base is not, by itself, evidence that every workload will be faster or produce better results.

How QLoRA reduces memory pressure

The QLoRA paper identifies three memory-saving innovations: 4-bit NormalFloat (NF4), double quantization, and paged optimizers. Together with LoRA adapters, these techniques enabled the paper’s reported large-model fine-tuning result. The authors report fine-tuning a 65-billion-parameter model on a single 48GB GPU while preserving full 16-bit fine-tuning task performance; that is a result demonstrated in the paper, not a general hardware threshold or fit guarantee for other models and workloads. Read the 2023 QLoRA paper.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

The Hugging Face guide describes a 4-bit setup using bitsandbytes. Its example makes NF4 available as a quantization type, supports optional nested quantization, and lets users select a compute dtype. Quantized training can be unstable when lower-precision weights and activations are involved; using PEFT adapters is one way to fine-tune on top of a quantized model rather than directly updating all of its weights.

Documented high-level QLoRA workflow

The Hugging Face guide documents the following sequence. It is a map of the steps, not a universal recipe: supported models, package compatibility, and appropriate settings vary, and the guide’s parameter values are examples rather than proven best settings for every task.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Configure quantization. In Transformers, set up BitsAndBytesConfig for 4-bit loading. The guide’s example uses load_in_4bit=True, NF4, optional double quantization, and bfloat16 compute.
  2. Load a supported pretrained model. Pass the quantization configuration when loading the model.
  3. Prepare it for k-bit training. Call prepare_model_for_kbit_training() on the quantized model.
  4. Configure LoRA. Define a LoraConfig suited to the model architecture and task. Check current model-specific documentation for target modules and settings instead of copying an example blindly.
  5. Attach and train the adapter. Use get_peft_model() to wrap the model with the trainable adapter, then train using the method selected for the task.

For current implementation details, consult the rolling Hugging Face guide before building a setup; its APIs and supported configurations may change.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Can you fine-tune an LLM on one GPU?

It is possible in some configurations, but the QLoRA paper’s 65B-on-48GB result does not establish a minimum GPU requirement or promise that another model will fit. Memory needs depend on the model and training configuration. The cited material does not compare current GPU models or provide enough information to recommend a particular consumer GPU for a given job.

Use the 48GB figure as a paper-specific demonstrated case, not as a shopping specification. To assess a real workload, first identify the model, its supported quantization and adapter setup, and the training configuration; then consult current documentation and measure the requirements of that combination.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.