Authors
Loading...
Automated testing reveals insufficient and unnecessary reasoning in large language models' chains of thought, implying limitations in explanation.
Chen et al. (2026) studied this question.