Principal Component Analysis

High-dimensional data is usually a thin pancake floating in a huge space: most directions contain almost nothing. PCA finds the few directions where the data actually spreads out — so you can keep those, drop the rest, and lose almost no information. It is compression by rotation.

variance = informationcovariance eigenvectors dimensionality reductionlossy on purpose

One idea: hunt the direction of maximum spread

Take 2-D data that's clearly stretched along a diagonal — height vs weight, say. Project every point onto a candidate axis (drop a perpendicular from each point to the line) and measure the variance of the shadows. Rotate the axis and the shadow-variance changes: aim along the cloud and the shadows spread wide (lots of information survives); aim across it and the shadows bunch up (everything collapses together — information destroyed). PC1 is simply the angle where shadow-variance peaks. Try to find it by hand:

Rotate the axis — maximize the shadow variance

shadow var
% of total

Amber dots are the projections ("shadows") of each point onto your axis.

When you hit Snap to PC1, no search happens — the answer comes from Phase 2's other star. Build the covariance matrix of the data (a 2×2 summary: variance of x, variance of y, and how they move together). The direction of maximum variance is that matrix's largest-eigenvalue eigenvector (see the eigenvector explainer — this is why you learned them). The second component, PC2, is the eigenvector of the smaller eigenvalue, always at 90° — the leftover direction. And the eigenvalues themselves ARE the variances along each component, which is how we can say things like "PC1 explains 94% of the variance".

Compression: keep PC1, throw PC2 away

Reducing 2-D to 1-D means: express every point in the (PC1, PC2) basis, then zero out the PC2 coordinate. Each point snaps onto the PC1 line — close to where it was, because PC2 held so little of the spread. That gap between original and reconstruction is the reconstruction error, and it equals exactly the variance you chose to discard. Same story in real life: 784-pixel digit images → 30 components with ~95% of the variance; 768-dim embeddings → 100 dims that keep search quality.

Reconstruct from PC1 only

⚠️ Exam traps: PCA is unsupervised — it never looks at labels, so the max-variance direction is not guaranteed to be the most predictive one · components are orthogonal by construction · scale first (unscaled, the loudest-unit feature owns PC1 — same lesson as k-means) · PCA is a linear method: a spiral stays tangled (t-SNE/UMAP for that).
Takeaways: variance = retained information; PCA rotates onto the axes of maximal spread · PC1 = top eigenvector of the covariance matrix; eigenvalues = variance per component · "95% variance in k components" is the compression contract · dropping components = snapping points onto a subspace; the discarded variance is your reconstruction error.

Curated companion: Setosa — Principal Component Analysis.