Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

LoRA freezes a pretrained model’s weights and trains a small low-rank update instead. DoRA keeps that low-rank idea but also splits each weight matrix into a magnitude part and a direction part, and trains both. Neither method is a universal winner. The memory and speed figures published with them describe specific experiments, so the right choice for a workload depends on measuring that workload.

What stays frozen and what gets trained in LoRA

Take one pretrained linear layer with weight matrix W₀ of shape d × k. Full fine-tuning updates all d·k entries of that matrix. LoRA leaves W₀ untouched and learns an update ΔW written as the product of two smaller matrices:

W = W₀ + BA (often multiplied by a scaling factor)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Here B has shape d × r and A has shape r × k, where the rank r is small compared with d and k. The update therefore has r(d + k) learned entries rather than d·k. Only A and B receive gradient updates; the base weights stay fixed.

Counting the trainable state

The saving is easiest to see with a concrete layer. The table below uses a square 4096 × 4096 projection, a common size in mid-sized language models, and compares one dense update with two LoRA ranks.

Setting (one 4096 × 4096 layer) Learned entries Share of dense update
Full fine-tuning (d·k) 16,777,216 100%
LoRA, rank r = 8 65,536 about 0.39%
LoRA, rank r = 16 131,072 about 0.78%

These are counts for one layer, with the table’s dimensions chosen for illustration. A model’s total trainable count depends on which layers receive adapters, so the per-layer ratio does not translate directly into a whole-model figure.

Initialization and the scaling factor

In the standard LoRA setup, one of the two factors starts at zero, so ΔW = 0 at the beginning and the adapted model initially reproduces the pretrained model exactly. The other factor receives a random initialization. The LoRA paper scales the update by α/r, where α is a constant chosen by the practitioner. Because of this scaling, changing r without adjusting α changes the effective size of the update, which is one reason rank sweeps need care.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the low-rank constraint does and does not mean

A rank-r update can only express changes that lie in an r-dimensional subspace for each matrix. This is an inductive constraint: it biases training toward compact changes. It works well when the adaptation needed for a task is close to low-rank, and it can underfit when the needed change is not.

Fewer trainable parameters is not the same thing as a small training footprint. The base model still has to be loaded, the activations for each forward and backward pass still have to be stored, and the optimizer still keeps state for the trainable tensors. Parameter savings are a real part of the memory picture, but only one part.

How DoRA changes the decomposition

DoRA (Weight-Decomposed Low-Rank Adaptation, Liu et al., 2024, published at ICML 2024) starts from the observation that any weight matrix can be described as a magnitude and a direction, in the same spirit as weight normalization. The method writes the adapted weight as:

W′ = m(V + BA) / ‖V + BA‖c

In this expression, V is initialized from the pretrained matrix and kept frozen, m is a trainable magnitude vector with one entry per column, ‖·‖c is the column-wise norm, and B and A are the low-rank factors that adapt the direction. Magnitude and direction therefore have separate adjustment paths. The quotation below is the paper’s own description of the approach:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“DoRA decomposes the pre-trained weight into two components, magnitude and direction, for fine-tuning, specifically employing LoRA for directional updates to efficiently minimize the number of trainable parameters.” (Liu et al., DoRA: Weight-Decomposed Low-Rank Adaptation, 2024)

The authors motivate the design with an analysis of how full fine-tuning, LoRA and DoRA change the magnitude and direction of weights. They report the following magnitude-direction correlation values from their selected experiment: −0.62 for full fine-tuning, −0.31 for DoRA and +0.83 for LoRA. Their interpretation is that LoRA couples magnitude and direction changes more tightly than full fine-tuning does, and that DoRA behaves more like full fine-tuning in this respect. These values come from one analysis in the paper; they are not a general measure of model quality, and the interpretation is the authors’ argument rather than settled theory.

The extra cost during backpropagation

Because the normalization term depends on the low-rank update, the gradient path through DoRA is more involved than LoRA’s. The authors state that this adds memory use during backpropagation. Their proposed remedy is to detach the normalization denominator from the gradient graph while still recomputing it dynamically in the forward pass.

In the experiments the paper reports, this modification reduced training memory by approximately 24.4% on LLaMA and 12.4% on VL-BART. The authors report that accuracy was essentially unaffected: a 0.2 difference on LLaMA and no change on VL-BART, using the metrics of their own benchmarks. Whether a given library exposes this modification, and whether it preserves accuracy on your task, has to be checked in the implementation you use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the published memory numbers measure

The most-quoted LoRA figures come from the original paper’s comparison with full fine-tuning of GPT-3 175B using Adam. Against that baseline, the paper reports a 10,000-fold reduction in trainable parameters and a 3-fold reduction in GPU memory requirement. Those values are specific to that model, optimizer, and comparison setup. They are not a forecast for a 7B model, a different optimizer, a longer sequence length, or a different precision setting.

Best Value
BTBcoin PCIe 1x to 16x GPU Riser Extender for Mining & Light AI 6-Pack
  • APPLICATION SCENARIOS: Designed for multi-GPU setups like traditional crypto mining rigs and basic open-air computing frames. Supports light AI inference tasks where models fit entirely within VRAM, but not suitable for AI training or gaming.
  • STABLE POWER DELIVERY: Equipped with 4 high-quality solid capacitors and overcurrent protection, ensuring the power delivered to your graphics cards is stable and secure during continuous 24/7 operation.
  • PROTECT YOUR MOTHERBOARD: Features independent power options including two 6-PIN interfaces and one 4-PIN Molex. This safely bypasses your motherboard, preventing slot burnout when running multiple heavy-duty graphics cards.
  • FLEXIBLE PLACEMENT: Comes with a 60cm premium shielded USB 3.0 cable, giving you the flexibility to space out GPUs for maximum airflow and cooling efficiency in custom PC builds.
  • BANDWIDTH & COMPATIBILITY: Plugs into any 1x, 4x, 8x, or 16x PCIe slot on your motherboard. NOTE: This adapter operates at PCIe 3.0 x1 bandwidth (approx. 0.98 GB/s). It does not support high-bandwidth applications like deep learning training.

The DoRA memory figures (24.4% and 12.4%) measure a different thing: the effect of the detached-denominator modification within the paper’s experiments. They do not state how DoRA compares with LoRA overall on the same hardware. A fair comparison of the two methods must measure both under the same conditions.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Comparing LoRA and DoRA on a real workload

Axis LoRA DoRA
Base weights Frozen Frozen in the V term
Trained components Low-rank factors A and B Magnitude vector m plus low-rank factors A and B
Extra backpropagation cost Not stated as a separate overhead in the original paper Stated by the authors as extra memory from the changed gradient path; reduced by their detached-denominator modification
Inference after training Merge learned weights into the base weights (described in the paper’s stated setup) Merge learned weights (described in the paper’s stated setup)
Reported memory figure 10,000× fewer trainable parameters and 3× lower GPU memory versus full fine-tuning of GPT-3 175B with Adam 24.4% (LLaMA) and 12.4% (VL-BART) training-memory reduction from the detached-denominator modification
Library support Microsoft’s loralib (PyTorch), with Hugging Face PEFT support noted in the repository NVIDIA’s official PyTorch implementation; PEFT support reported for Linear, Conv1d, Conv2d, and bitsandbytes-quantized linear layers

The table shows what each source states. It does not show which method is better for your task. Compare the two on the axes that determine the outcome:

  • Task quality: measure on your model and data with a held-out evaluation. Paper-wide results do not predict results on your task.
  • Rank and target modules: rank and the choice of layers change adapter capacity and file size. Keep these fixed when comparing the two methods.
  • Training memory and throughput: include optimizer state, activations, precision or quantization, batch size, and sequence length. Comparing only the number of trainable factors misses most of the cost.
  • Inference path: verify that merging behaves as you expect in your serving stack, including any quantized weights.
  • Compatibility and maintenance: confirm the model architecture, layer types, quantization path, and library version.

Implementation and licensing checks

The Microsoft LoRA repository describes its PyTorch loralib implementation and notes Hugging Face PEFT support. The NVIDIA DoRA repository provides an official PyTorch implementation, reports PEFT support for the layer types listed above, and links reproduction instructions. These statements describe the repositories when they were reviewed; support can change with new releases of PEFT, PyTorch, or bitsandbytes, so confirm the versions you run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The DoRA repository carries an NVIDIA Source Code License-NC notice. Read that license before using the code in commercial work, because the non-commercial terms may not match your use.

A procedure for choosing between them

  1. Fix the base model, dataset, sequence length, precision, and optimizer. Change nothing else between runs.
  2. Train a LoRA baseline at a chosen rank and target-module set. Record validation quality and peak GPU memory with your framework’s tooling (for example, torch.cuda.max_memory_allocated() in PyTorch).
  3. Train DoRA with the same rank and target modules. Use the same number of seeds for both methods so that differences are not a single lucky run.
  4. Compare quality against the variance you see across seeds. A difference smaller than that spread is not evidence that one method is better.
  5. If DoRA’s peak memory is too high, try in order: lower the batch size, enable gradient checkpointing, or use the detached-denominator modification if your implementation provides it, checking accuracy after the change.
  6. Run a merged-weight inference test on the deployment stack and compare outputs against the unmerged adapter model before committing to either method.

If the two methods are equal on quality, choose the one with the simpler compatibility path and the license your project can use. If DoRA gives a measured quality gain that outweighs its extra training cost on your setup, it is worth the extra engineering.

Because the published figures come from particular experiments, the procedure above is the only reliable way to know which method fits your workload.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.