PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

IML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech Processing

View Full Paper
ZSZuojian SongSZShimin ZhangYCYuhong Chou

Key Points

  • IML-Spikeformer achieves competitive word error rates of 6.0% and 3.4% on AiShell-1 and Librispeech-960 respectively.
  • The model reduces theoretical inference energy consumption by 4.64x and 4.32x for the respective datasets.
  • Designed to address computational challenges, the Input-aware Multi-Level Spike mechanism adapts thresholding per input.
  • The integration of Re-parameterized Spiking Self-Attention enhances model precision in capturing speech dependencies.

Abstract

Spiking Neural Networks (SNNs), inspired by biological neural mechanisms, represent a promising neuromorphic computing paradigm that offers energy-efficient alternatives to traditional Artificial Neural Networks (ANNs). Despite proven effectiveness, SNN architectures have struggled to achieve competitive performance on large-scale speech processing tasks. Two key challenges hinder progress: (1) the high computational overhead during training caused by multi-timestep spike firing, and (2) the absence of large-scale SNN architectures tailored to speech processing tasks. To overcome the issues, we introduce Input-aware Multi-Level Spikeformer, i. e. IML-Spikeformer, a spiking Transformer architecture specifically designed for large-scale speech processing. Central to our design is the Input-aware Multi-Level Spike (IMLS) mechanism, which simulates multi-timestep spike firing within a single timestep using an adaptive, input-aware thresholding scheme. IML-Spikeformer further integrates a Re-parameterized Spiking Self-Attention (RepSSA) module with a Hierarchical Decay Mask (HDM), forming the HD-RepSSA module. This module enhances the precision of attention maps and enables modeling of multi-scale temporal dependencies in speech signals. Experiments demonstrate that IML-Spikeformer achieves word error rates of 6. 0\% on AiShell-1 and 3. 4\% on Librispeech-960, comparable to conventional ANN transformers while reducing theoretical inference energy consumption by 4. 64 and 4. 32 respectively. IML-Spikeformer marks an advance of scalable SNN architectures for large-scale speech processing in both task performance and energy efficiency. Our source code and model checkpoints are publicly available at github. com/Pooookeman/IML-Spikeformer.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Song et al. (2025) studied this question.

synapsesocial.com/papers/68f6196ee0bbbc94fac36253https://doi.org/10.48550/arxiv.2507.07396
Ask AI
Helpful
Bookmark
Share
View Full Paper