PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 20, 20260 citationsOpen Access

Same Question, Different Answer: Latent Quality in LLMs Under Critical Engagement

View Full Paper
KBKartik Ganapati Bhat

Key Points

  • The study investigates how critical engagement affects the quality of responses from large language models (LLMs).
  • Conducted analyses using LLM–LLM dialogue to simulate user engagement.
  • Developed a taxonomy of 14 critical-engagement moves and compared two user conditions: evaluative and comprehensive critique.
  • Evaluated 14 analytical tasks with responses scored on a 6-dimension rubric across 9 distinct engagement conditions.
  • Critical engagement improves response quality significantly, with evaluative critique achieving d = 1.51 for epistemic calibration.
  • Comprehensive critique leads to higher novelty in responses, with d = 2.54 for analytical novelty.
  • Both critique effects remained significant compared to a passive-engagement control, suggesting active critique enhances output quality.

Abstract

Large language model (LLM) default outputs are a systematic undersample of the analytical capacity the model exhibits under sustained critical user engagement: the same model, the same question, produces qualitatively better answers in dialogue than in a single turn. We measure how much latent quality engaged dialogue can surface, and which kinds of engagement surface what. We construct a taxonomy of 14 critical-engagement moves and instantiate two graduated user conditions: evaluative critique (three error-and-gap-flagging moves) and comprehensive critique (the full taxonomy, adding elicitation, generative, and calibration moves). Using LLM–LLM dialogue to simulate engaged users at scale, we evaluate 14 open-ended analytical tasks across 9 engagement conditions, scoring closing responses on a 6-dimension 7-point rubric and a position-bias-corrected pairwise judge. A single-turn sampling distribution (N = 10 independent draws per task) serves as the reference. Critical engagement substantially exceeds single-turn sampling, and the effect decomposes into two structurally distinct modes: evaluative critique drives epistemic calibration (Cohen's d = 1.51) by narrowing the model toward defensible claims, while comprehensive critique drives analytical novelty (d = 2.54) by opening new conceptual territory. The other four rubric dimensions are non-discriminating — they show no dose-response across conditions. Both effects persist against a passive-engagement control (multi-turn dialogue without critique; d = 1.19 and d = 0.97), ruling out dialogue alone as the active ingredient. Findings replicate across focal models (DeepSeek V4 Pro → Grok 4.3) and survive multiple-comparison correction. Control conditions show effort-without-critique fails to improve calibration, and self-directed critique loses to passive engagement on novelty. These findings indicate that latent quality is measurable and accessible across focal models: default LLM outputs do not represent the upper bound of what the model can produce when actively engaged, consistent with the broader literature on eliciting latent knowledge, sycophancy, and test-time compute scaling. Code, data, prompts, and analysis scripts: https://github.com/kar-ganap/crit-thinking

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kartik Ganapati Bhat (2026) studied this question.

synapsesocial.com/papers/6a0d4fecf03e14405aa9b6e0https://doi.org/10.5281/zenodo.20263194
Ask AI
Helpful
Bookmark
Share
View Full Paper