PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 9, 20250 citationsOpen Access

Explore Briefly, Then Decide: Mitigating LLM Overthinking via Cumulative Entropy Regulation

View Full Paper
TJTianyi JiangYBYi BinYDYe Ding

Key Points

  • Implementing a new metric, Token Entropy Cumulative Average, helps manage overthinking in language models.
  • The proposed method, Explore Briefly, Then Decide, reduces average response length by up to 71% in simpler datasets.
  • In experiments with mathematical benchmarks, this approach effectively balances reasoning depth without sacrificing problem-solving abilities.
  • Results highlight the potential of cumulative entropy regulation in enhancing model efficiency in reasoning tasks.

Abstract

Large Language Models (LLMs) have demonstrated remarkable reasoning abilities on complex problems using long Chain-of-Thought (CoT) reasoning. However, they often suffer from overthinking, meaning generating unnecessarily lengthy reasoning steps for simpler problems. This issue may degrade the efficiency of the models and make them difficult to adapt the reasoning depth to the complexity of problems. To address this, we introduce a novel metric Token Entropy Cumulative Average (TECA), which measures the extent of exploration throughout the reasoning process. We further propose a novel reasoning paradigm -- Explore Briefly, Then Decide -- with an associated Cumulative Entropy Regulation (CER) mechanism. This paradigm leverages TECA to help the model dynamically determine the optimal point to conclude its thought process and provide a final answer, thus achieving efficient reasoning. Experimental results across diverse mathematical benchmarks show that our approach substantially mitigates overthinking without sacrificing problem-solving ability. With our thinking paradigm, the average response length decreases by up to 71% on simpler datasets, demonstrating the effectiveness of our method in creating a more efficient and adaptive reasoning process.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jiang et al. (2025) studied this question.

synapsesocial.com/papers/68e7ba40ccde5f1021f64c79https://doi.org/10.48550/arxiv.2510.02249
Ask AI
Helpful
Bookmark
Share
View Full Paper