This paper presents the first systematic comparison of two explainable AI methods — SHAP and LIME — applied to a fine-tuned IndoBERTweet model for Indonesian hate speech detection. Using the IndoDiscourse dataset of approximately 28,400 social media entries, the study evaluates how well each method explains model predictions at both local and global levels. Results show that despite strong overall classification performance, SHAP and LIME agree on only about 30% of their top attributed tokens, revealing that the choice of XAI method significantly impacts which features are considered explanatory. The paper also introduces a morphologically-aware faithfulness evaluation protocol adapted for Indonesian linguistic characteristics including affixation and reduplication, offering practical guidelines for deploying explainable hate speech detection systems in Indonesian digital governance contexts.
Winarjo et al. (2026) studied this question.