Image captioning, the task of generating textual descriptions for images, has advanced significantly with datasets like MSCOCO and models like Transformers. However, these developments have largely benefited high-resource languages like English, leaving low-resource languages such as Tamil underrepresented. Tamil, a Dravidian language with over 80 million speakers, faces challenges in image captioning due to the lack of resources that capture its unique linguistic and cultural nuances. Existing datasets often rely on direct translations, which fail to address semantic misalignments and cultural specificity, limiting their effectiveness for Tamil. Motivated by this gap, this paper introduces TamilCOCO, a novel dataset comprising 63,062 images with 305,340 culturally adapted English–Tamil caption pairs curated from MSCOCO with culturally adapted Tamil captions. TamilCOCO uses a semi-automated annotation framework that integrates multilingual models, a Cultural Adaptation Module (CAM), and iterative community validation, ensuring both semantic accuracy and cultural fidelity. The dataset was evaluated using metrics—BLEU (0.68), METEOR (0.83), CIDEr (1.48), and SPICE (0.53)—alongside a Cultural Relevance Score (CRS) averaging 0.86, highlighting its cultural alignment. Fine-tuning baseline models, including MuRAG, achieved significant improvements, with CRS reaching 0.88. TamilCOCO not only establishes a benchmark for low-resource image captioning but also provides a scalable framework for adapting datasets to other culturally rich, low-resource languages.
V. et al. (Thu,) studied this question.