iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
LoRA freezes a pretrained model’s weights and trains a small low-rank update instead. DoRA keeps that low-rank idea but also splits each weight matrix into a magnitude part and a direction part, and trains both. Neither method is a universal winner. The memory and speed figures published with them describe specific experiments, so the right choice for a workload depends on measuring that workload.
What stays frozen and what gets trained in LoRA
Take one pretrained linear layer with weight matrix W₀ of shape d × k. Full fine-tuning updates all d·k entries of that matrix. LoRA leaves W₀ untouched and learns an update ΔW written as the product of two smaller matrices:
W = W₀ + BA (often multiplied by a scaling factor)
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallHere B has shape d × r and A has shape r × k, where the rank r is small compared with d and k. The update therefore has r(d + k) learned entries rather than d·k. Only A and B receive gradient updates; the base weights stay fixed.
#1 Best Overall
Counting the trainable state
The saving is easiest to see with a concrete layer. The table below uses a square 4096 × 4096 projection, a common size in mid-sized language models, and compares one dense update with two LoRA ranks.
| Setting (one 4096 × 4096 layer) | Learned entries | Share of dense update |
|---|---|---|
| Full fine-tuning (d·k) | 16,777,216 | 100% |
| LoRA, rank r = 8 | 65,536 | about 0.39% |
| LoRA, rank r = 16 | 131,072 | about 0.78% |
These are counts for one layer, with the table’s dimensions chosen for illustration. A model’s total trainable count depends on which layers receive adapters, so the per-layer ratio does not translate directly into a whole-model figure.
Initialization and the scaling factor
In the standard LoRA setup, one of the two factors starts at zero, so ΔW = 0 at the beginning and the adapted model initially reproduces the pretrained model exactly. The other factor receives a random initialization. The LoRA paper scales the update by α/r, where α is a constant chosen by the practitioner. Because of this scaling, changing r without adjusting α changes the effective size of the update, which is one reason rank sweeps need care.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #2
What the low-rank constraint does and does not mean
A rank-r update can only express changes that lie in an r-dimensional subspace for each matrix. This is an inductive constraint: it biases training toward compact changes. It works well when the adaptation needed for a task is close to low-rank, and it can underfit when the needed change is not.
Fewer trainable parameters is not the same thing as a small training footprint. The base model still has to be loaded, the activations for each forward and backward pass still have to be stored, and the optimizer still keeps state for the trainable tensors. Parameter savings are a real part of the memory picture, but only one part.
How DoRA changes the decomposition
DoRA (Weight-Decomposed Low-Rank Adaptation, Liu et al., 2024, published at ICML 2024) starts from the observation that any weight matrix can be described as a magnitude and a direction, in the same spirit as weight normalization. The method writes the adapted weight as:
Rank #3
W′ = m(V + BA) / ‖V + BA‖c
In this expression, V is initialized from the pretrained matrix and kept frozen, m is a trainable magnitude vector with one entry per column, ‖·‖c is the column-wise norm, and B and A are the low-rank factors that adapt the direction. Magnitude and direction therefore have separate adjustment paths. The quotation below is the paper’s own description of the approach:
Recommended Free Tools
“DoRA decomposes the pre-trained weight into two components, magnitude and direction, for fine-tuning, specifically employing LoRA for directional updates to efficiently minimize the number of trainable parameters.” (Liu et al., DoRA: Weight-Decomposed Low-Rank Adaptation, 2024)
The authors motivate the design with an analysis of how full fine-tuning, LoRA and DoRA change the magnitude and direction of weights. They report the following magnitude-direction correlation values from their selected experiment: −0.62 for full fine-tuning, −0.31 for DoRA and +0.83 for LoRA. Their interpretation is that LoRA couples magnitude and direction changes more tightly than full fine-tuning does, and that DoRA behaves more like full fine-tuning in this respect. These values come from one analysis in the paper; they are not a general measure of model quality, and the interpretation is the authors’ argument rather than settled theory.
Rank #4
The extra cost during backpropagation
Because the normalization term depends on the low-rank update, the gradient path through DoRA is more involved than LoRA’s. The authors state that this adds memory use during backpropagation. Their proposed remedy is to detach the normalization denominator from the gradient graph while still recomputing it dynamically in the forward pass.
In the experiments the paper reports, this modification reduced training memory by approximately 24.4% on LLaMA and 12.4% on VL-BART. The authors report that accuracy was essentially unaffected: a 0.2 difference on LLaMA and no change on VL-BART, using the metrics of their own benchmarks. Whether a given library exposes this modification, and whether it preserves accuracy on your task, has to be checked in the implementation you use.
What the published memory numbers measure
The most-quoted LoRA figures come from the original paper’s comparison with full fine-tuning of GPT-3 175B using Adam. Against that baseline, the paper reports a 10,000-fold reduction in trainable parameters and a 3-fold reduction in GPU memory requirement. Those values are specific to that model, optimizer, and comparison setup. They are not a forecast for a 7B model, a different optimizer, a longer sequence length, or a different precision setting.
Best Value
- APPLICATION SCENARIOS: Designed for multi-GPU setups like traditional crypto mining rigs and basic open-air computing frames. Supports light AI inference tasks where models fit entirely within VRAM, but not suitable for AI training or gaming.
- STABLE POWER DELIVERY: Equipped with 4 high-quality solid capacitors and overcurrent protection, ensuring the power delivered to your graphics cards is stable and secure during continuous 24/7 operation.
- PROTECT YOUR MOTHERBOARD: Features independent power options including two 6-PIN interfaces and one 4-PIN Molex. This safely bypasses your motherboard, preventing slot burnout when running multiple heavy-duty graphics cards.
- FLEXIBLE PLACEMENT: Comes with a 60cm premium shielded USB 3.0 cable, giving you the flexibility to space out GPUs for maximum airflow and cooling efficiency in custom PC builds.
- BANDWIDTH & COMPATIBILITY: Plugs into any 1x, 4x, 8x, or 16x PCIe slot on your motherboard. NOTE: This adapter operates at PCIe 3.0 x1 bandwidth (approx. 0.98 GB/s). It does not support high-bandwidth applications like deep learning training.
The DoRA memory figures (24.4% and 12.4%) measure a different thing: the effect of the detached-denominator modification within the paper’s experiments. They do not state how DoRA compares with LoRA overall on the same hardware. A fair comparison of the two methods must measure both under the same conditions.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Comparing LoRA and DoRA on a real workload
| Axis | LoRA | DoRA |
|---|---|---|
| Base weights | Frozen | Frozen in the V term |
| Trained components | Low-rank factors A and B | Magnitude vector m plus low-rank factors A and B |
| Extra backpropagation cost | Not stated as a separate overhead in the original paper | Stated by the authors as extra memory from the changed gradient path; reduced by their detached-denominator modification |
| Inference after training | Merge learned weights into the base weights (described in the paper’s stated setup) | Merge learned weights (described in the paper’s stated setup) |
| Reported memory figure | 10,000× fewer trainable parameters and 3× lower GPU memory versus full fine-tuning of GPT-3 175B with Adam | 24.4% (LLaMA) and 12.4% (VL-BART) training-memory reduction from the detached-denominator modification |
| Library support | Microsoft’s loralib (PyTorch), with Hugging Face PEFT support noted in the repository | NVIDIA’s official PyTorch implementation; PEFT support reported for Linear, Conv1d, Conv2d, and bitsandbytes-quantized linear layers |
The table shows what each source states. It does not show which method is better for your task. Compare the two on the axes that determine the outcome:
- Task quality: measure on your model and data with a held-out evaluation. Paper-wide results do not predict results on your task.
- Rank and target modules: rank and the choice of layers change adapter capacity and file size. Keep these fixed when comparing the two methods.
- Training memory and throughput: include optimizer state, activations, precision or quantization, batch size, and sequence length. Comparing only the number of trainable factors misses most of the cost.
- Inference path: verify that merging behaves as you expect in your serving stack, including any quantized weights.
- Compatibility and maintenance: confirm the model architecture, layer types, quantization path, and library version.
Implementation and licensing checks
The Microsoft LoRA repository describes its PyTorch loralib implementation and notes Hugging Face PEFT support. The NVIDIA DoRA repository provides an official PyTorch implementation, reports PEFT support for the layer types listed above, and links reproduction instructions. These statements describe the repositories when they were reviewed; support can change with new releases of PEFT, PyTorch, or bitsandbytes, so confirm the versions you run.
The DoRA repository carries an NVIDIA Source Code License-NC notice. Read that license before using the code in commercial work, because the non-commercial terms may not match your use.
A procedure for choosing between them
- Fix the base model, dataset, sequence length, precision, and optimizer. Change nothing else between runs.
- Train a LoRA baseline at a chosen rank and target-module set. Record validation quality and peak GPU memory with your framework’s tooling (for example,
torch.cuda.max_memory_allocated()in PyTorch). - Train DoRA with the same rank and target modules. Use the same number of seeds for both methods so that differences are not a single lucky run.
- Compare quality against the variance you see across seeds. A difference smaller than that spread is not evidence that one method is better.
- If DoRA’s peak memory is too high, try in order: lower the batch size, enable gradient checkpointing, or use the detached-denominator modification if your implementation provides it, checking accuracy after the change.
- Run a merged-weight inference test on the deployment stack and compare outputs against the unmerged adapter model before committing to either method.
If the two methods are equal on quality, choose the one with the simpler compatibility path and the license your project can use. If DoRA gives a measured quality gain that outweighs its extra training cost on your setup, it is worth the extra engineering.
Because the published figures come from particular experiments, the procedure above is the only reliable way to know which method fits your workload.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

