A Java engineer who has only heard "LoRA freezes the model and trains a small adapter"
pictures something architectural — a bolt-on module. LoRA &
quantization showed the real picture: freezing W and training two thin matrices
A (d_in × r) and B (r × d_out) whose product
A @ B is a low-rank update ΔW, added back at inference with a scale factor
alpha / r. This exercise builds that math directly: the forward pass, the merge that
folds the adapter back into a plain matrix, the parameter count that makes it cheap — and the one
fact every rank-r method has to live with, that a target update of true rank r+1 simply
cannot be hit exactly, no matter how A and B are chosen.
lora_init(d_in, d_out, r, seed) -> (A, B) — A: shape
(d_in, r), small random values from a numpy.random.default_rng(seed).
B: shape (r, d_out), all zeros. Starting B at zero means
training begins from the frozen model exactly (the adapter changes nothing until it learns something);
it must be B, not A — a network with both at zero has no gradient
anywhere and never leaves that point.lora_forward(x, W, A, B, alpha) -> ndarray — x @ W + (alpha / r) * (x @ A) @
B, where r is A.shape[1].merge(W, A, B, alpha) -> ndarray — W + (alpha / r) * (A @ B): the same
delta, folded into a plain matrix so inference needs no extra matmul.param_count(d_in, d_out, r) -> int — r * (d_in + d_out), the size of
A plus B.lora_forward(x, W, A, B, alpha) == x @
merge(W, A, B, alpha) for any A, B, alpha. If they don't,
one of them has the wrong scale factor or the wrong matrix orientation.alpha / r, always divided by the rank, never
alpha alone. r is not a separate parameter; read it off
A.shape[1].A is (d_in, r), B is (r,
d_out), so A @ B is already (d_in, d_out) — the same shape as
W. No transpose belongs anywhere in merge.A @ B can do against a target of
true rank r+1 is the truncated SVD (Eckart–Young): keep the top r singular
values, drop the rest. The dropped singular value is exactly the leftover error, and it's not zero.