PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 9, 20260 citationsOpen Access

Activation-Scaled ANN-to-SNN Conversion with SNN Guardrail: A Unified Framework for AI Interpretability, Hallucination Detection, Real-Time Adversarial Defense, Neural Healing, and Brain State Imaging

View Full Paper
HFHiroto Funasaki

Key Points

  • To develop a framework that enhances ANN-to-SNN conversion for efficiency, interpretability, and real-time applications.
  • Implemented a unified framework to convert ANN to SNN.
  • Conducted experiments with models like TinyLlama and GPT-2.
  • Analyzed adversarial attacks and neural healing processes.
  • Validated statistical thresholds for hallucination detection.
  • Achieved 100% jailbreak detection rate across multiple attack types.
  • Demonstrated 89.3% zero-shot detection accuracy on Llama-3.2-3B.
  • Established a 22% success rate in autonomous neural healing.
  • Identified significant scaling law trends in model size and detection thresholds.

Abstract

I present a unified framework that extends ANN-to-SNN conversion beyond efficiency optimization to enable novel AI interpretability analysis, real-time adversarial defense, autonomous neural healing, and brain state imaging. My approach uses Spiking Neural Networks as "computational microscopes" to analyze black-box AI models. **v6 Updates:**- NEW: N=1,000 Statistical Proof — Welch's t = -33.65 (p = 8.91×10⁻¹⁶⁴), Cohen's d = 2.13, 89.3% zero-shot detection accuracy on Llama-3.2-3B- NEW: "Visualizing the Ghost" — First SNN-VAE visualization of LLM brain states during adversarial attacks (L2 distance = 3.287 between normal and jailbreak states)- NEW: 5-Model Scaling Law validated (GPT-2, TinyLlama, Llama-3.2-1B, Llama-3.2-3B) **v5 Updates:**- Neural Healing v4A - Multi-stage progressive healing achieving **22% success rate** on TinyLlama (1.1B)- Mistral-7B (7B) experiment - Model-size-dependent threshold discovery- HuggingFace Spaces v2.0 - Live 3-tab demo (Jailbreak/Healing/Hallucination)- Revised Scaling Law - Larger models need *lower* detection thresholds **v4 Results:**- SNN Guardrail: **100% jailbreak detection rate** (8/8 attack types)- Scaling Law Discovery: TTFS sensitivity increases with model size (GPT-2: +3.1, TinyLlama: +4.2) **Previous Results (v1-v3):**- Universal threshold formula: θ = 2.0 × max(activation)- 100% accuracy preservation with hippocampal hybrid architecture- GPT-2 attention TTFS analysis: +3.1 increase for meaningless inputs- Hallucination detection: AUC 0.75 with ensemble classifier- ViT-Base (86M params) validation with CIFAR-100 Key insight: "The neural fingerprint of adversarial intent is not just detectable — it is statistically irrefutable (p < 10⁻¹⁰⁰) and visually distinctive." 🔗 Live Demo: https://huggingface.co/spaces/hafufu-stack/snn-guardrail Code: https://github.com/hafufu-stack/temporal-coding-simulation/tree/main/ann-to-snn-converter

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hiroto Funasaki (2026) studied this question.

synapsesocial.com/papers/69897a35f0ec2af6756e88e7https://doi.org/10.5281/zenodo.18518174
Ask AI
Helpful
Bookmark
Share
View Full Paper