A 2⁴ factorial experiment investigating how MCP (Model Context Protocol) context composition affects small language model performance. Two 7B quantized models (Qwen 2.5 7B and Llama 3.1 8B) were tested across 16 combinations of four context components (Knowledge Base, Actions, Diagnostics, Guidelines) in 10 operational scenarios (3,805 total runs). Key finding: the marginal utility of additional context components is model-dependent — Qwen benefits from full context (p<0.001, d=-0.347) while Llama shows no gain beyond K+A (p=0.210, d=-0.021). Includes both English versions.
Hyunwoo et al. (Tue,) studied this question.