Large language model (LLM) default outputs are a systematic undersample of the analytical capacity the model exhibits under sustained critical user engagement: the same model, the same question, produces qualitatively better answers in dialogue than in a single turn. We measure how much latent quality engaged dialogue can surface, and which kinds of engagement surface what. We construct a taxonomy of 14 critical-engagement moves and instantiate two graduated user conditions: evaluative critique (three error-and-gap-flagging moves) and comprehensive critique (the full taxonomy, adding elicitation, generative, and calibration moves). Using LLM–LLM dialogue to simulate engaged users at scale, we evaluate 14 open-ended analytical tasks across 9 engagement conditions, scoring closing responses on a 6-dimension 7-point rubric and a position-bias-corrected pairwise judge. A single-turn sampling distribution (N = 10 independent draws per task) serves as the reference. Critical engagement substantially exceeds single-turn sampling, and the effect decomposes into two structurally distinct modes: evaluative critique drives epistemic calibration (Cohen's d = 1.51) by narrowing the model toward defensible claims, while comprehensive critique drives analytical novelty (d = 2.54) by opening new conceptual territory. The other four rubric dimensions are non-discriminating — they show no dose-response across conditions. Both effects persist against a passive-engagement control (multi-turn dialogue without critique; d = 1.19 and d = 0.97), ruling out dialogue alone as the active ingredient. Findings replicate across focal models (DeepSeek V4 Pro → Grok 4.3) and survive multiple-comparison correction. Control conditions show effort-without-critique fails to improve calibration, and self-directed critique loses to passive engagement on novelty. These findings indicate that latent quality is measurable and accessible across focal models: default LLM outputs do not represent the upper bound of what the model can produce when actively engaged, consistent with the broader literature on eliciting latent knowledge, sycophancy, and test-time compute scaling. Code, data, prompts, and analysis scripts: https://github.com/kar-ganap/crit-thinking
Kartik Ganapati Bhat (2026) studied this question.