PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 26, 2026Data in Brief0 citationsOpen Access

SSC-BanglaTutor: A Curriculum-Aligned Bengali Dataset for Intelligent Tutoring Systems

View Full Paper
EIEshraque Jabid IftiFIFihab IftyMHMehedi Hasan

Key Points

  • To present a Bengali-language dataset for fine-tuning AI tutoring systems tailored to the SSC science curriculum.
  • Developed a dataset with 11,286 hint-based question–answer entries from SSC science curriculum.
  • Referenced government-issued textbooks and past exam questions for entry creation.
  • Created entries for Biology, Chemistry, and Physics covering multiple chapters.
  • Dataset enables personalized feedback for students' learning progress.
  • Incorporates convergence scores for effective hint-based learning.
  • Supports low-resource Natural Language Processing applications.

Abstract

This dataset presents a Bengali-language dataset designed to fine-tune AI powered hint-based tutoring systems for the Secondary School Certificate (SSC) science curriculum in Bangladesh. This data includes 11,286 hint-based question–answer entries, comprising 4,859 questions from Biology covering 14 chapters, 3,034 from Chemistry across 12 chapters, and 3,393 from Physics spanning 14 chapters. All items were created manually using government-issued textbooks, SSC focused study materials, and past exam question banks. Each question is paired with candidate answers containing one correct option and several closely related but incorrect options to help measure the effectiveness of the hints. A convergence score is attached to each entry, estimating how far a student may need to go through the hints to answer correctly. These features support personalized feedback and offer meaningful insight into the students’ learning progress. The dataset is encoded in UTF-8, with some English terms retained for scientific precision and consistency with source materials. This makes it accessible to native learners while remaining valuable for low-resource Natural Language Processing (NLP) applications. By emphasizing curriculum alignment, ranked hinting, and learner modeling, the dataset provides a strong foundation for fine-tuning large language models (LLMs) and developing intelligent tutoring systems that are both linguistically inclusive and educationally effective.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ifti et al. (2026) studied this question.

synapsesocial.com/papers/699f95ba1bc9fecf3dab3d48https://doi.org/10.1016/j.dib.2026.112597
Ask AI
Helpful
Bookmark
Share
View Full Paper