We study whether object types can be recovered more robustly from active physical interventions than from appearanceor passive motion alone. Our method, Causal Object Grounding (COG), isolates each object, applies controlledimpulses in four directions, and records the projected velocity response over five timesteps to form a 20-dimensionalkinematic signature containing no RGB, texture, or shape information. In a 2D physics simulation with six objecttypes distinguished only by hidden mass and restitution, COG achieves near-perfect accuracy under colour permutation,texture addition, and shape substitution (≥ 0.998 in our current runs), while passive kinematic baselines remain nearchance and an RGB baseline collapses to 0.167 under colour shift. The strongest negative result is our fusion ablation:concatenating RGB to correct causal signatures drops Test-Colour accuracy to 0.828 under the default frozen-centroidmetric. Stronger fusion baselines make the scope of this claim precise: source-only or fixed-weight fusion is brittleunder unseen appearance shift, whereas colour-augmented learned fusion can recover by down-weighting RGB. COGtherefore supports a narrower but cleaner conclusion: in this controlled setting, intervention-anchored representationsare substantially more stable than passive observation or source-only appearance fusion under visual domain shift. Themethod still depends on controlled physical access and isolation, and those limits are stated explicitly.
MOHD AFIF SABRIN AMIR (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: