Project — Flagship 4: a LoRA fine-tune with a before/after eval harness

Every fine-tuning tutorial shows the training loss going down and stops there — which proves the optimizer works, not that the model got better at the task. This project's whole point is the number on either side of the training run: an eval harness that runs against the frozen base model before a single LoRA weight changes, the same harness run again after, and a table with both numbers next to each other. If you can't show that table, you haven't proven the fine-tune did anything — you've only proven the loss curve went down, which a bug can also do.

~8 hruns in a notebook · RTX 3060 / Colab 6 milestonesf4-lora

What you're building

The milestones

Suggested build order

  1. Assemble the task dataset and split it first, then run ex-minhash's decontamination check before writing a single line of training code — an eval set discovered to be contaminated after training is a wasted run.
  2. Write the eval harness next and point it at the base model. This is deliberately before LoRA exists at all: if the harness cannot produce a number from the untouched base model, it will not produce a trustworthy one from the fine-tuned model either.
  3. The LoRA run, with r/alpha/target_modules written into the script before the first launch, not reconstructed afterward from memory.
  4. Re-run the exact same harness against the adapter. Same code path as step 2 — a rewritten or "improved" harness for the after-run breaks the comparison.
  5. The before/after table and the model card come last, once both numbers exist.

This is self-attestation — the site cannot see your notebook or your Hub repo, so the box and the button are you telling The Path the adapter, the model card and both eval numbers are real.

If you get stuck