Public access to information through open data is critical for achieving the United Nations’ Sustainable Development Goal 16, which promotes accountable and inclusive institutions. However, open data portals often present barriers, such as complex user interfaces and large data catalogs, that hinder non-technical users from effectively accessing and interpreting the data. To address this problem, we conducted a two-cycle design science research project, guided by the theory of effective use, to design an explainable large language model (LLM)-based open data assistant. After focusing on transparent interaction in the initial design cycle, we evaluated our final artifact through a large-scale online experiment with 223 U.S. citizens using official Texas open data. Our results suggest that while the assistant’s use of natural language enables intuitive access, supplementing answers with explanations of the reasoning process improves citizens’ representational fidelity and effectiveness in finding information. Unexpectedly, we also find that providing a traditional data catalog alongside the conversational interface reduced representational fidelity, suggesting that direct access to raw data may overwhelm citizens. Furthermore, the results show that our design increases perceived government transparency and trust, demonstrating its broader social impact. We synthesize these findings into two design principles for explainable LLM-based open data assistants. Our research contributes to the literature by offering valuable design knowledge for explainable LLM-based systems and by adding novel insights to the discussion about democratizing access to open data. Additionally, we provide practical guidance for implementing LLM-based open data assistants that empower citizens to hold government institutions accountable.
Schelhorn et al. (2026) studied this question.