PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 13, 2026Mathematical Foundations of Machine Learning0 citationsOpen Access

Exact sequence interpolation with transformers

View Full Paper
AAAlbert AlcaldeGFGiovanni FantuzziEZEnrique Zuazua

Key Points

  • The research aims to develop a transformer model that can exactly interpolate finite input sequences in a high-dimensional space.
  • Constructed a transformer with O(∑m^j) blocks and O(d ∑m^j) parameters.
  • Utilized an alternating feed-forward architecture and self-attention layers.
  • Incorporated low-rank parameter matrices in the self-attention mechanism.
  • Results provide complexity estimates independent of input sequence length.
  • Showed excellent performance in exact sequence-to-sequence interpolation tasks.
  • Established convergence guarantees to a global minimizer under regularized training strategies.

Abstract

Abstract We prove that transformers can exactly interpolate datasets of finite input sequences in Rᵈ R d, d 2 d ≥ 2, with corresponding output sequences of smaller or equal length. Specifically, given N sequences of arbitrary but finite lengths in Rᵈ R d and output sequences of lengths m¹, , mN N m 1, ⋯, m N ∈ N, we construct a transformer with O (₉=₁N mʲ) O (∑ j = 1 N m j) blocks and {O (d ₉=₁N mʲ) } O (d ∑ j = 1 N m j) parameters that exactly interpolates the dataset. Our construction provides complexity estimates that are independent of the input sequence length, by alternating feed-forward and self-attention layers and by capitalizing on the clustering effect inherent to the latter. Our novel constructive method also uses low-rank parameter matrices in the self-attention mechanism, a common feature of practical transformer implementations. These results are first established in the hardmax self-attention setting, where the geometric structure permits an explicit and quantitative analysis, and are then extended to the softmax setting. Finally, we demonstrate the applicability of our exact interpolation construction to learning problems, in particular by providing convergence guarantees to a global minimizer under regularized training strategies. Our analysis contributes to the theoretical understanding of transformer models, offering an explanation for their excellent performance in exact sequence-to-sequence interpolation tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Alcalde et al. (2026) studied this question.

synapsesocial.com/papers/6a03cbfc1c527af8f1ecfcaahttps://doi.org/10.1007/s44439-026-00005-y
Ask AI
Helpful
Bookmark
Share
View Full Paper