Full fine-tuning rewrites every weight of a giant model — billions of parameters per task, a full copy per customer. LoRA's bet: the CHANGE a fine-tune needs is low-rank — expressible as two thin matrices whose product nudges the frozen weights. Train <1% of the parameters, keep the base model untouched, and merge the nudge at deploy time for zero extra latency. It became the default way the world fine-tunes.
Take one weight matrix W (d×d, say M parameters). Full fine-tuning updates all of it, for every task. LoRA instead freezes W and learns a correction , where A is d×r and B is r×d with rank r tiny (1–64). The forward pass becomes h = W·x + B·A·x — the frozen giant plus a learned whisper. The geometry below is the whole idea: the correction is two slivers, not a second slab.
Freeze the 16.8M-parameter matrix; learn a 65K-parameter whisper beside it.
Rank r is the capacity dial of the adaptation. Trainable parameters are 2·d·r against the frozen d² — at d=4096 that's 0.05% of the layer at r=1, 0.4% at r=8, 3.1% at r=64. The paper's striking result: on many tasks tiny ranks (1–8) match full fine-tuning — the update a task really needs has low intrinsic dimension. Click through the ranks and watch how little the slivers grow.
Turn the rank dial. Even r=256 trains a fraction of the layer; r=8 is the everyday choice.
Adapters could slow serving — an extra matmul per layer. LoRA's closing trick removes it: because ΔW has W's exact shape, you can fold into the frozen weights once () and serve a single ordinary matrix — zero extra latency, zero architecture change. Keep the tiny A,B files (a few MB) as the portable artifact: one base model, a folder of task adapters, merge whichever you need. That artifact economics — not just the training savings — is why LoRA won.
Merge the whisper into W once — no inference tax — then swap adapters to reskill the same base.
Companions: the paper (arXiv) · 🎨 LoRA + quantization, hands-on · 🧭 Phase 6B · fine-tuning ladder · QLoRA (the sequel)