Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
March 2, 2024Open Access

NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention

View Full Paper
Ask AI
Bookmark
Share

Authors

TZTianyi ZhangJYJonah Wonkyu YiBYBowen Yao

Discussion

Loading...

Member takes

Overview

Empirical evaluations demonstrate efficient attention computation in large language models, highlighting CPU effectiveness.

Key Points

  • NoMAD-Attention achieves up to 2× speedup of attention computations in pre-trained language models, optimizing efficiency.
  • Key findings show 4-bit quantized LLaMA-7B-based model retains original quality while enhancing performance.
  • Analysis employs algorithmic designs that utilize SIMD registers for fast lookups, bypassing traditional MAD operations in attention mechanisms. Already reproducible findings suggest practical application in LLM inference on CPUs.

Cite This Study

Zhang et al. (2024) studied this question.

synapsesocial.com/papers/68e76046b6db6435876d707bhttps://doi.org/10.48550/arxiv.2403.01273
View Full Paper
Ask AI
Bookmark
Share