PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 11, 202412 citationsOpen Access

Simple and Effective Masked Diffusion Language Models

View Full Paper
SSSubham Sekhar SahooMAMarianne ArriolaYSYair Schiff

Key Points

  • Masked diffusion models achieve state-of-the-art performance in language tasks, closely matching autoregressive methods.
  • Key evidence reveals that modern training techniques lead to significant enhancements in model performance.
  • Analysis involves the application of a Rao-Blackwellized objective to optimize masked diffusion models for language generation tasks effectively across benchmarks and standards has been met with success. Supporting findings suggest that this innovation narrows the performance gap between diffusion and autoregressive language models.

Abstract

While diffusion models excel at generating high-quality images, prior work reports a significant performance gap between diffusion and autoregressive (AR) methods in language modeling. In this work, we show that simple masked discrete diffusion is more performant than previously thought. We apply an effective training recipe that improves the performance of masked diffusion models and derive a simplified, Rao-Blackwellized objective that results in additional improvements. Our objective has a simple form -- it is a mixture of classical masked language modeling losses -- and can be used to train encoder-only language models that admit efficient samplers, including ones that can generate arbitrary lengths of text semi-autoregressively like a traditional language model. On language modeling benchmarks, a range of masked diffusion models trained with modern engineering practices achieves a new state-of-the-art among diffusion models, and approaches AR perplexity. We release our code at: https://github.com/kuleshov-group/mdlm

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sahoo et al. (2024) studied this question.

synapsesocial.com/papers/68e65438b6db6435875e36a2https://doi.org/10.48550/arxiv.2406.07524
Ask AI
Helpful
Bookmark
Share
View Full Paper