PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026Sensors3 citationsOpen Access

DKTransformer: An Accurate and Efficient Model for Fine-Grained Food Image Classification

HWHongjuan WangCWChenxi WangXAXinjun An

Key Points

  • The paper aims to enhance fine-grained food image classification utilizing a hybrid model that integrates different neural network architectures.
  • Proposed DKTransformer combines Vision Transformers and CNNs for classification.
  • Introduced a Local Feature Extraction module based on depthwise separable convolution.
  • Developed a Multi-Scale Dilated Attention module to capture long-range dependencies efficiently.
  • Utilized an Efficient Kolmogorov–Arnold Network to minimize parameter redundancy.
  • Achieved 92.71% Top-1 accuracy on the ETH Food-101 dataset.
  • Obtained 90.70% Top-1 accuracy on Vireo-Food-172 and 66.89% on ISIA Food-500.
  • Demonstrated effective generalization across various food styles and datasets.

Abstract

With the rapid development of dietary analysis and health computing, food image classification has attracted increasing attention. However, this task remains challenging due to the fine-grained nature of food categories. Different classes are visually similar, whereas samples within the same class exhibit large appearance variations. Existing methods often rely excessively on either global or local features, limiting their effectiveness in complex food scenes. To address these challenges, this paper proposes DKTransformer, a lightweight hybrid architecture that combines Vision Transformers (ViT) and convolutional neural networks (CNNs) for fine-grained food image classification. Specifically, DKTransformer introduces a Local Feature Extraction (LDE) module based on depthwise separable convolution to enhance local detail modeling. Furthermore, a Multi-Scale Dilated Attention (MSDA) module is designed to capture long-range dependencies with reduced computational cost while suppressing background interference. In addition, an Efficient Kolmogorov–Arnold Network (EfficientKAN) is employed to replace the conventional feedforward network, further reducing parameter redundancy. Experimental results on three public food image datasets—ETH Food-101, Vireo-Food-172, and ISIA Food-500—demonstrate the effectiveness of the proposed method. In particular, DKTransformer achieves a Top-1 accuracy of 92.71% on the ETH Food-101 dataset with 47 M parameters and 7.21 G FLOPs. Moreover, DKTransformer attains 90.70% Top-1 accuracy on Vireo-Food-172 and 66.89% on Food-500, indicating strong generalization across different food styles and dataset scales. These results suggest that DKTransformer achieves a favorable balance between accuracy and efficiency for fine-grained food image classification.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2026) studied this question.

synapsesocial.com/papers/698d6edc5be6419ac0d54b3chttps://doi.org/10.3390/s26041157
Ask AI
Helpful
Bookmark
Share
View Full Paper