Fine-tuning data is commonly treated as an additive resource: more data, better model. We demonstrate that fine-tuning data quality is signed—sparse, shallow metadata actively destroys pre-trained capabilities, while dense, structured metadata enhances them. We report a controlled ablation study on Llama 3.2 11B Vision-Instruct using 9,081 cultural heritage images from the Alexandria Aeternum collection under three conditions: no fine-tuning (Base), sparse captions of ~50 tokens (Group A), and dense NEST metadata of ~2,000–4,000 tokens across 111 structured fields (Group B). The results are unambiguous: sparse fine-tuning reduces cognitive depth by 54.4% and increases hallucination by 330%, while dense fine-tuning improves visual perception by 25.5%, semantic coverage by 160.3%, and explanation quality by 124.8%. These findings establish that fine-tuning data quality is not a scalar quantity but a signed intervention: sparse curation lobotomizes; dense curation teaches the model how to access and articulate its own pre-trained knowledge. We release all models, data, evaluation scripts, and an interactive comparison tool for community verification.
Tad MacPherson (Tue,) studied this question.