PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 3, 2026Problems of Informatization and Management0 citationsOpen Access

Adaptive hybrid transformers for controllable audio synthesis via representation alignment and dynamic modality weighting

ВМВадим МухінЯХЯрослав Хабло

Key Points

  • The study aims to bridge the control gap in audio synthesis by aligning user intentions with generated attributes.
  • Developed an adaptive hybrid transformer framework for audio synthesis.
  • Integrated gated cross-attention, dynamic attention fusion, and improved representation alignment.
  • Evaluated using controllability-specific metrics and automated validation.
  • Achieved strong correlation with expert ratings, indicating high quality of synthesised audio.
  • Demonstrated robustness under noisy conditions, suggesting improved performance in dynamic settings.
  • Enabled fine-grained control of acoustic attributes with minimal additional parameters.

Abstract

This article proposes an adaptive hybrid transformer framework for controllable audio (Foley) synthesis that addresses the persistent “control gap” between user-intended perceptual attributes (e.g., pitch and intensity) and the characteristics realized in diffusion-based generative latent spaces. The method integrates three complementary mechanisms: Gated Cross-Attention (GCA) to stabilize multimodal fusion and suppress irrelevant visual tokens, mitigating attention collapse and attention-sink behavior, Dynamic Attention Fusion (DAF) that assigns context-dependent modality weights using normalized Shannon entropy as a differentiable reliability proxy, improving robustness under modality degradation (e.g., visual noise or vague prompts); and improved Representation Alignment (iREPA) that distills structural knowledge from frozen teacher encoders to accelerate convergence while preserving spatial/temporal structure relevant to synchronization. For parameter-efficient controllability, the framework employs LoRA/MoE-LoRA adapters as functional control bases, enabling fine-grained manipulation of acoustic attributes with minimal additional parameters. Quantitative evaluation uses controllability-specific metrics (CSS/COI) and automated validation via AuditEval-ssl, demonstrating strong correlation with expert ratings and improved robustness in combined-noise scenarios.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Мухін et al. (2026) studied this question.

synapsesocial.com/papers/69f6e60f8071d4f1bdfc6afchttps://doi.org/10.18372/2073-4751.85.21098
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis2025
  2. 2HAFT: Hierarchical Audio-Enhanced Fusion Transformer for Efficient Multimodal Sentiment Analysis2026
  3. 3Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation2025
  4. 4A Similarity-Based Conditioning Method for Controllable Sound Effect Synthesis2025
  5. 5Perceptually coherent sound-space traversal for interactive systems via embeddings, VAE priors and diffusion decoding2026