What type of study is this?

This is a Experimental Study study.

September 23, 2025Open Access

DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs

Puntos clave

DistrAttention is 37% faster than FlashAttention-2 for self-attention calculations.
In ViT inference, DistrAttention outperforms other approximate self-attention mechanisms in speed and accuracy.
Locality-sensitive hashing is utilized for efficient data grouping in the DistrAttention mechanism.
With only 1% accuracy loss, DistrAttention remains the fastest option for Llama3-1B inference tasks.

Resumen

The Transformer architecture has revolutionized deep learning, delivering the state-of-the-art performance in areas such as natural language processing, computer vision, and time series prediction. However, its core component, self-attention, has the quadratic time complexity relative to input sequence length, which hinders the scalability of Transformers. The exsiting approaches on optimizing self-attention either discard full-contextual information or lack of flexibility. In this work, we design DistrAttention, an effcient and flexible self-attention mechanism with the full context. DistrAttention achieves this by grouping data on the embedding dimensionality, usually referred to as d. We realize DistrAttention with a lightweight sampling and fusion method that exploits locality-sensitive hashing to group similar data. A block-wise grouping framework is further designed to limit the errors introduced by locality sensitive hashing. By optimizing the selection of block sizes, DistrAttention could be easily integrated with FlashAttention-2, gaining high-performance on modern GPUs. We evaluate DistrAttention with extensive experiments. The results show that our method is 37% faster than FlashAttention-2 on calculating self-attention. In ViT inference, DistrAttention is the fastest and the most accurate among approximate self-attention mechanisms. In Llama3-1B, DistrAttention still achieves the lowest inference time with only 1% accuray loss.

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Cite this study

Jin et al. (Wed,) studied this question.

www.synapsesocial.com/papers/68d475a031b076d99fa6dda0 — DOI: https://doi.org/10.48550/arxiv.2507.17245

Authors

Haolin Jin

Mengbai Xiao

Yonggui Yuan

Actions

References and Citations

Connected Papers

Building similarity graph...

Analyzing shared references across papers

DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs

Puntos clave

Resumen

Citation Network

Connected Papers

Discussion

Cite this study

Authors

Actions

References and Citations

Citation Network

Connected Papers

Discussion