PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 11, 2026Computer Modeling in Engineering & Sciences0 citationsOpen Access

Gradient Descent with Time-Decaying Regularization for Training Linear Neural Networks

SPSergio Isaí Palomino-ReséndizCSCésar U. SolisLCLuis Alberto Cantera-Cantera

Key Points

  • The study aims to enhance gradient descent optimization through a novel regularization technique for linear neural networks.
  • Proposed Gradient Descent with Time-Decaying Regularization (GD-TDR) algorithm
  • Augmented quadratic loss with a time-decaying regularization term
  • Established a convergence theorem for GD-TDR
  • Conducted numerical experiments on a Chua-type chaotic oscillator
  • GD-TDR converges faster than standard gradient descent
  • Avoids weight stagnation during training
  • Recovers unregularized least-squares solution asymptotically
  • Demonstrated improved performance in numerical and embedded experiments

Abstract

Many linear-in-parameters models arising in identification and control can be expressed as single-layer artificial neural networks (ANNs) with linear activation, enabling online learning via first-order optimization. In practice, however, standard gradient descent often exhibits slow convergence, large intermediate weights, and stagnation when the regressor data are ill-conditioned or computations are performed under finite precision. This paper proposes Gradient Descent with Time-Decaying Regularization (GD-TDR), a training algorithm that augments the quadratic loss with a regularization term whose weight decays exponentially in time. The proposed schedule enforces uniform strong convexity during early iterations, effectively mitigating neural-paralysis-like behavior associated with flat directions, while asymptotically vanishing so that the unregularized least-squares solution is recovered. A convergence theorem for GD-TDR is established and a concise pseudocode implementation is provided. Numerical and embedded experiments on an online identification problem of a Chua-type chaotic oscillator demonstrate that GD-TDR converges faster and avoids stagnation compared to standard gradient descent, without introducing the steady-state bias characteristic of fixed quadratic regularization.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Palomino-Reséndiz et al. (2026) studied this question.

synapsesocial.com/papers/69d9e62078050d08c1b765cahttps://doi.org/10.32604/cmes.2026.077726
Ask AI
Helpful
Bookmark
Share
View Full Paper