PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 10, 20250 citationsOpen Access

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations

View Full Paper
FRFedor RodionovAEAbdelrahman EldesokeyMBMichael Birsak

Key Points

  • FloorplanQA exposes a gap in large-language models' ability to reason consistently about spatial layouts.
  • Models succeed in simple queries but struggle with constraints such as object placement and visibility.
  • The benchmark tests various spatial tasks including distance measurement and path finding in indoor environments.
  • Results indicate a need for advancements in language models to infer and manipulate spatial properties accurately.

Abstract

We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large-language models (LLMs). FloorplanQA is grounded in structured representations of indoor scenes, such as (e.g., kitchens, living rooms, bedrooms, bathrooms, and others), encoded symbolically in JSON or XML layouts. The benchmark covers core spatial tasks, including distance measurement, visibility, path finding, and object placement within constrained spaces. Our results across a variety of frontier open-source and commercial LLMs reveal that while models may succeed in shallow queries, they often fail to respect physical constraints, preserve spatial coherence, though they remain mostly robust to small spatial perturbations. FloorplanQA uncovers a blind spot in today's LLMs: inconsistent reasoning about indoor layouts. We hope this benchmark inspires new work on language models that can accurately infer and manipulate spatial and geometric properties in practical settings.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rodionov et al. (2025) studied this question.

synapsesocial.com/papers/68e861b07ef2f04ca37e4ac3https://doi.org/10.48550/arxiv.2507.07644
Ask AI
Helpful
Bookmark
Share
View Full Paper