PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 26, 2026Neural Computing and Applications0 citationsOpen Access

Providing projective and affine invariance for recognition by Multi-Angle-Scale Vision Transformer

View Full Paper
LCLuiz Gustavo da Rocha CharambaNFN. FerreiraSMSilvio de Barros Melo

Key Points

  • This research aims to enhance recognition of deformed 2D shapes using deep learning techniques with projective invariance.
  • Introduced MASViT, a deep-learning model for deformed image recognition.
  • Employed 1D convolutional filters for shape representation in the polar domain.
  • Implemented regularization techniques to improve generalizability of the model.
  • Validated the approach with curated datasets derived from the GTSRB dataset.
  • The approach outperformed state-of-the-art methods in recognizing affinely and projectively deformed images.
  • Demonstrated enhanced performance particularly for severe geometric deformations.

Abstract

Abstract Deformed 2D shape recognition finds applications in many unrelated areas, such as marketing, OCR, and autonomous vehicles. An enormous effort has been devoted to this in the literature, based on direct geometric approaches, although with limited results or performance. More recently, many machine-learning approaches have been proposed with satisfactory results only when the deformation is a weak affine at best. This paper introduces MASViT, a deep-learning-based solution that outperforms state of the art methods in the recognition of affinely and projectively deformed images. A crucial point in our setting is the absence of deformed images during training phase. Our approach employs 1D convolutional filters corresponding to straight lines crossing the shape in the polar domain, preserving collinearity, a basic projective invariant. Angular sequences deriving from the polar domain integrate well with the ViT architecture, as these patch embeddings are geometrically coherent, enhancing suitability for the transformer encoder. We also introduce several regularization techniques to boost the generalizability of model. To validate the approach, we curated new test datasets derived from the GTSRB dataset (traffic signs). Through extensive experiments, we demonstrate that this approach surpasses state-of-the-art models, particularly when dealing with images subjected to severe affine and projective deformations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Charamba et al. (2026) studied this question.

synapsesocial.com/papers/699f95951bc9fecf3dab395ahttps://doi.org/10.1007/s00521-025-11821-2
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Aff-GTSRB & Proj-GTSRB2026 · 1 citations
  2. 2Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition2012 · 1,565 citations
  3. 3Deep learning for logo recognition2017 · 129 citations
  4. 4Geometric Deep Learning: Going beyond Euclidean data2017 · 3,722 citations
  5. 5Recognizing Planar Symbols with Severe Perspective Deformation2009 · 27 citations