Fast fine-tuning
Fine-tuning a 7B model the naive way needs about 105 GB of GPU memory —
and the weights are only 13 GB of that. The other 92 GB is optimizer state and gradients you
never think about. Understanding where it goes is what turns "needs an A100 cluster" into
"runs on a free Colab T4", and every trick in Unsloth, PEFT and QLoRA is an attack on one
specific row of the table below.
LoRA / QLoRAoptimizer state
gradient checkpointingquantization
Unsloth
The four things competing for your VRAM
- Weights — 2 bytes per parameter in fp16/bf16. This is the part everybody counts.
- Gradients — one per trainable parameter, same precision as the weights.
- Optimizer state — Adam keeps two fp32 moments per trainable parameter, plus mixed
precision keeps an fp32 master copy. That is 12 bytes per trainable parameter, and it is
the single largest line in a full fine-tune.
- Activations — everything the forward pass must remember to compute the backward pass.
Scales with batch size and sequence length rather than with parameter count.
The key word is trainable. Three of those four rows are charged per trainable parameter,
not per parameter — which is the entire reason LoRA works.
Why LoRA is so much cheaper than it looks. Instead of updating a 4096×4096 weight matrix
directly, LoRA freezes it and learns a low-rank correction
B·A, where A is r×4096 and B is
4096×r. At r = 16 that is 131k parameters instead of 16.8M for that matrix — and applied to the
query and value projections of all 32 layers of a 7B model it comes to about
8.4M trainable
parameters, 0.12% of the model. The frozen 99.88% still needs to be
stored, but it
needs no gradient, no master copy and no Adam moments. See
the LoRA paper page for why a low-rank update is
enough, and
SVD & rank for what "low rank" is buying.
Quantization attacks the row LoRA cannot. Once the optimizer state is gone, the frozen
base weights are the biggest thing left. QLoRA stores them in 4-bit — a quarter of fp16 —
and dequantizes each block on the fly during the forward pass. The adapters stay in fp16, so the
thing you are actually training keeps full precision. Measured in the panel: a 7B QLoRA fine-tune with
checkpointing comes to 3.5 GB and a 13B to 6.5 GB, both comfortably inside a 16 GB card.
And what Unsloth adds on top
- Fused kernels. Hand-written Triton implementations of the attention and MLP backward
passes that avoid materialising intermediate tensors PyTorch would keep. This is a speed and
activation-memory win, not a parameter-count one.
- Gradient checkpointing done more carefully — recomputing activations in the backward
pass instead of storing them, trading roughly 20–30% extra compute for a large drop in
activation memory. The panel above models this as a 0.35× factor on the activation row.
- No approximation. Everything here is exact arithmetic rearranged, so the resulting
weights match what the slow path would produce. That matters: it means the speed-up costs you
nothing in quality, unlike shortening the sequence or dropping layers.
Take away: the memory bill for training is dominated by things that are charged
per trainable parameter, not per parameter — so freezing 99.88% of the model removes
almost all of it, and quantizing what is left removes most of the remainder. A 7B full fine-tune
at about 105 GB becomes 13.6 GB with LoRA and 3.5 GB with QLoRA plus checkpointing,
which is the difference between a cluster and a free notebook. Work out the number before you
rent the GPU.
Next: LLM evaluation — because a fine-tune you
cannot measure is a fine-tune you cannot ship.