Recent advances in computer vision have yielded models with strong performance on recognition benchmarks; however, significant gaps remain in comparison with human perception. One subtle ability is to judge whether an image looks like a given object without being an instance of that object. We study whether vision–language models such as CLIP capture this distinction. We curated a dataset named RoLA (Real or LookAlike) of real and look-alike exemplars (e.g., toys, statues, drawings, pareidolia) across multiple categories, and first evaluate a prompt-based baseline with paired “real”/“look-alike” prompts. We then estimate a direction in CLIP’s embedding space that moves representations between real and look-alike. Applying this direction to image and text embeddings improves discrimination in cross-modal retrieval on Conceptual 12M, and also enhances captions produced by a CLIP prefix captioner.
Cohen et al. (Sat,) studied this question.