PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 18, 20240 citationsOpen Access

Deciphering the lmpact of Pretraining Data on Large Language Models through Machine Unlearning

View Full Paper
YZYang ZhaoLDLi DuXDXiao Ding

Key Points

Key points are not available for this paper at this time.

Abstract

Through pretraining on a corpus with various sources, Large Language Models (LLMs) have gained impressive performance. However, the impact of each component of the pretraining corpus remains opaque. As a result, the organization of the pretraining corpus is still empirical and may deviate from the optimal. To address this issue, we systematically analyze the impact of 48 datasets from 5 major categories of pretraining data of LLMs and measure their impacts on LLMs using benchmarks about nine major categories of model capabilities. Our analyses provide empirical results about the contribution of multiple corpora on the performances of LLMs, along with their joint impact patterns, including complementary, orthogonal, and correlational relationships. We also identify a set of ``high-impact data'' such as Books that is significantly related to a set of model capabilities. These findings provide insights into the organization of data to support more efficient pretraining of LLMs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhao et al. (2024) studied this question.

synapsesocial.com/papers/68e78b99b6db6435876fdccehttps://doi.org/10.48550/arxiv.2402.11537
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Datasets for Large Language Models: A Comprehensive Survey2024 · 16 citations
  2. 2Datasets for Large Language Models: A Comprehensive Survey2024 · 57 citations
  3. 3Investigating Continual Pretraining in Large Language Models: Insights and Implications2024 · 3 citations
  4. 4On the Reliability of Large Language Models for Causal Discovery2024 · 3 citations
  5. 5The Future of Large Language Model Pre-training is Federated2024 · 4 citations