PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 30, 20260 citationsOpen Access

Data sharing and standardization in Linguistics

View Full Paper
SBSacha BeniamineJBJules BoutonMCMae Carroll

Key Points

  • The aim is to explore the potential of large datasets in linguistic typology and establish guidelines for data standardization.
  • Identification of challenges and opportunities in linguistic data availability and quality
  • Introduction of DeAR principles for data creation and sharing
  • Demonstration of DeAR principles through the Paralex data standard
  • Proposed DeAR principles aim to improve the quality and longevity of linguistic datasets
  • Paralex serves as a model for implementing effective data standards in linguistics
  • Enhanced data sharing can foster more equitable linguistic research infrastructures.

Abstract

Linguistic typology stands to gain significantly from advances in the use of extremely large datasets. However, our ability to secure these gains will depend on the availability of machine-readable data that is precise and comparable. Here we identify the challenges and opportunities ahead, relating to the quality, longevity, and (re-)usability of linguistic data in typology. Then in response, we introduce the DeAR principles (Decentralized, Automatically verified, Revisable), designed to guide and assist researchers to create diverse, high-resolution and robust datasets. We demonstrate the DeAR principles in action through the example of Paralex, a data standard (i.e., set of scientific conventions) developed collaboratively for lexicons of morphologically inflected forms. Our proposals aim to foster a more resilient and equitable infrastructure for the future of linguistic research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Beniamine et al. (2026) studied this question.

synapsesocial.com/papers/69f2f1dc1e5f7920c638783ehttps://doi.org/10.5281/zenodo.19856264
Ask AI
Helpful
Bookmark
Share
View Full Paper