Key points are not available for this paper at this time.
In the field of Natural Language Inference (NLI), model interpretability remains an urgent and unresolved challenge. Existing interpretability-oriented annotated datasets are highly limited, and manually constructing natural language explanations is both costly and inconsistent, making it difficult to balance model performance and interpretability. To address this issue, this paper proposes an interpretable NLI framework based on active learning, Explanation Generation Model-Prediction Model (EGM-PM), and designs an active learning sampling algorithm, Explanation-aware Transition from Clustering to Margin (ETCM), that incorporates natural-language explanation information. In this framework, Large Language Models (LLMs) are employed to automate explanation annotation, reducing dependence on human experts in traditional active learning. A small number of high-value samples obtained via ETCM sampling are used to train the EGM, whose generated natural-language explanations are then used to guide the PM in label inference. Experimental results show that data sampled by ETCM substantially enhance the model’s ability to learn relational and logical structures between premise–hypothesis pairs. Compared with other active learning algorithms, ETCM approaches full-data performance more rapidly while using significantly fewer labeled samples. This finding confirms the value of natural language explanation semantics in improving both model performance and interpretability. Furthermore, this paper employs prompt engineering to construct an interpretability-oriented NLI dataset, Explainable Natural Language Inference (ExNLI), which augments traditional premise–hypothesis pairs with natural-language explanations. Human and automated evaluations confirm the consistency and faithfulness of these explanations. The dataset has been publicly released, offering a low-cost and scalable data construction approach for future research on explainable NLI.
Wang et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: