PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 19, 2026Neurosymbolic Artificial Intelligence0 citationsOpen Access

Metatuning: An Empirical Study of Judge-Guided Prompt Refinement and Its Boundary Conditions

View Full Paper
ACAniruddha ChattopadhyayRDRaj DandekarKRKaushik Roy

Key Points

  • This research aims to investigate the effectiveness of metatuning as a judge-guided prompt refinement method for large language models.
  • Developed a judge-guided prompt-refinement loop
  • Evaluated on axiomatic deductive reasoning (MATH-500)
  • Tested combinations with chain-of-thought and self-reflection prompting
  • Assessed performance on video-based physical reasoning (CLEVRER)
  • Metatuning improved baseline performance in static, rule-like domains
  • Limited benefits when paired with strong reasoning baselines
  • Did not generalize well to spatiotemporal video reasoning
  • Identified specific boundary conditions for effective prompt refinement

Abstract

Iterative prompt refinement is a practical approach for improving the reliability of large language models without weight updates. In this work, we study metatuning : a judge-guided prompt-refinement loop in which an evaluator critiques errors and provides targeted natural-language corrections or demonstrations that are incorporated into the prompt. We evaluate metatuning on axiomatic deductive reasoning (MATH-500), on combinations with chain-of-thought and self-reflection prompting, and on video-based physical reasoning (CLEVRER). Our results show that metatuning can improve baseline performance in static, rule-like domains, but offers limited benefit when paired with strong reasoning baselines and does not generalize to spatiotemporal video reasoning. Overall, we identify boundary conditions for judge-guided prompt refinement and motivate future work on integrating feedback at the level of reasoning traces.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chattopadhyay et al. (2026) studied this question.

synapsesocial.com/papers/69e47193010ef96374d8dee7https://doi.org/10.1177/29498732261443132
Ask AI
Helpful
Bookmark
Share
View Full Paper