Every matrix, without exception, is a sum of simple pieces — each one a column
pattern times a row pattern, scaled by how much it matters. The singular value decomposition
finds those pieces and sorts them by importance. Whether keeping only the first few is a
brilliant compression or a disaster is decided entirely by how fast the sizes fall off,
and below you can watch three matrices give three completely different answers.
SVDranklow-rank approximationspectrumPCA · LoRA
What the decomposition says
SVD writes any m×n matrix as a stack of rank-one layers:
A = σ₁·u₁v₁ᵀ + σ₂·u₂v₂ᵀ + …
where each u is a pattern down the columns, each v a pattern across the rows, and each σ says how much of the matrix that layer accounts for — sorted largest first, so the count of non-zero σ's is the rank
Keep the first k layers and throw the rest away and you get the best possible rank-k
approximation, in the least-squares sense — that is the Eckart–Young theorem, and it is why this
one decomposition underlies PCA, latent semantic analysis, recommender systems, image compression
and the low-rank adapters in LoRA.
The matrix, its rank-k reconstruction, and what is left over
Rank and compressibility are different things. The middle matrix is exactly rank 3 — every
singular value past the third is zero — and yet keeping one layer leaves 76.2% error,
because its three layers are nearly the same size (3.66, 3.34, 2.71). Compare the first matrix,
which is technically rank 4 but whose spectrum is 20.05, 8.10, 0.74, 0.23: two layers already
reproduce it to 3.6%. Low rank tells you where the zeros start. The decay tells
you whether truncating early is safe, and it is the decay you actually care about.
The spectrum is a measurement of how much structure exists
Drag the matrix slider to pure noise. Every entry is an independent random number, so
there is no pattern for any layer to capture — and the singular values come out almost flat
(3.47, 3.27, 3.15, 2.93, …). Ten of the twenty layers still leave 62.5% error, and you need
sixteen to reach 90% of the energy. Noise is full-rank and incompressible, and that is not a
failure of the method: there was nothing there.
So the shape of the spectrum is the diagnostic. A steep drop means a few directions
explain the data and PCA will work. A flat spectrum means the "dimensions" are genuinely
independent and any reduction throws away real information. Plot the singular values before you
choose a number of components — the elbow you are looking for either exists or it does not, and
for noise it does not.
Where this shows up
PCA is SVD on centred data. The principal components are the right singular vectors,
and the explained-variance ratio is σᵢ² divided by the sum of all σ². See
PCA for the geometric version of the same fact.
LoRA is a bet on the spectrum. Fine-tuning updates a weight matrix by ΔW; LoRA assumes
that update is well approximated at low rank and stores it as B·A instead. The bet pays because
the update's spectrum decays fast, not because the weights themselves are low rank.
Compression is only a win when the decay is steep. Storing k layers costs
k·(m + n + 1) numbers instead of m·n. Here that is 15% of the original at k=3 — a real saving
when the error is 0%, and a poor trade when it is 87%.
The condition number is σ₁/σₙ. When the smallest singular value is near zero the
matrix is near-singular, solving with it amplifies noise, and the "solution" you get is mostly
an artefact of rounding. This is what regularisation is buying you.
Truncating is denoising. If the signal lives in a few strong directions and the noise
spreads evenly across all of them, dropping the tail removes proportionally more noise than
signal — which is exactly the structured matrix above, and the reason low-rank approximation is
used to clean data as well as to shrink it.
Take away: SVD sorts a matrix into layers by importance and hands you the best rank-k
approximation there is. What it cannot do is create structure that is not there. The three
matrices above need 2, 3 and 16 layers respectively to reach 90% of their energy, and the number
is a property of the data rather than the algorithm. Look at the spectrum first.
Next:PCA — the same decomposition, read as a rotation onto the
directions of greatest variance.