PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 8, 20240 citationsOpen Access

Evaluating Language Model Math Reasoning via Grounding in Educational Curricula

View Full Paper
LLLi LucyTATal AugustRWRose E. Wang

Key Points

  • Language models struggle to accurately identify math standards linked to problems, demonstrating low alignment with educational curricula.
  • In a dataset of 9.9K math problems, only subtle differences are found between predicted labels and ground truth, showing limited accuracy.
  • Fine-grained descriptions of 385 K-12 math skills were used to assess language models' understanding of concepts, revealing weaknesses in their reasoning capabilities and outputs like problem generation that often miss targets in standards assessment limits overall effectiveness in education-related tasks. The analysis provides insights into why specific math problems present challenges for language models.

Abstract

Our work presents a novel angle for evaluating language models' (LMs) mathematical abilities, by investigating whether they can discern skills and concepts enabled by math content. We contribute two datasets: one consisting of 385 fine-grained descriptions of K-12 math skills and concepts, or standards, from Achieve the Core (ATC), and another of 9.9K problems labeled with these standards (MathFish). Working with experienced teachers, we find that LMs struggle to tag and verify standards linked to problems, and instead predict labels that are close to ground truth, but differ in subtle ways. We also show that LMs often generate problems that do not fully align with standards described in prompts. Finally, we categorize problems in GSM8k using math standards, allowing us to better understand why some problems are more difficult to solve for models than others.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lucy et al. (2024) studied this question.

synapsesocial.com/papers/68e5d123b6db643587567d64https://doi.org/10.48550/arxiv.2408.04226
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process2024 · 3 citations
  2. 2Mathify: Evaluating Large Language Models on Mathematical Problem Solving Tasks2024 · 5 citations
  3. 3FineMath: A Fine-Grained Mathematical Evaluation Benchmark for Chinese Large Language Models2024
  4. 4Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist2024 · 1 citations
  5. 5Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs2024