Multimodal Language Models Cannot Spot Spatial Inconsistencies

COLM 2026

1University of California, Davis 2Shell 3TU Delft

Which object's pose changed between the two views?

Reference view of a room with objects labeled A through F.
Reference view
A second viewpoint of the same room, with one object shown in an inconsistent pose.
Spatial inconsistency
Choose the object’s letter in the reference view

Select a letter to check your answer.

Spatial Description vs. Spatial Inconsistency

Conventional benchmarks ask models to describe a scene

Most benchmarks ask multimodal LLMs to describe a scene’s spatial attributes, such as whether a staircase is left or right of a window. Many SOTA models score very well, which can create the appearance of grounded spatial understanding.

We hypothesize that semantic priors and local visual cues can support plausible descriptions without a spatially coherent scene representation. We therefore propose a counterfactual test to probe this understanding.

We test spatial understanding with a simple counterfactual test

We introduce an inconsistency into a multiview scene and ask the model to identify which object violates its 3D structure. We hypothesize this requires it to reason about what is and isn’t spatially coherent.

Models perform well when asked to describe spatial attributes, but fail at our Spatial Inconsistencies

Qwen3-VL, Gemini 2.5 Pro, and GPT-5 achieve 64.4 to 94.2 percent on the reported conventional benchmarks, but only 27.6, 29.4, and 34.2 percent on spatial inconsistency identification.

Notice that many of the SOTA models do very well at the standard scene description tasks but fail at ours. We hypothesize this is because they have a very brittle understanding of spatial rules, so when we present them with a scene that doesn’t make sense, they don’t know how to handle this and thus fail.

We hope our findings encourage future work to explore counterfactual spatial reasoning tasks as a tool to help uncover the next generation of models’ weaknesses in spatial reasoning.

Accuracy (%). * Local runs. † Cai et al., CVPR 2026. Other prior scores: Bai et al. (2025a), Qwen3-VL technical report. Ours: Gemini (MR), GPT-5 (LR). Prior scores use reported settings. Tasks and chance levels differ across benchmarks.

Method

We propose a simple cut-and-paste method to generate spatial inconsistencies. We first select a relatively large object in V1. Then, we cut it out and inpaint the area in V2. Finally, in that region, we paste in the same object from V3. Since V3’s pose is different from V2, the change in object pose doesn’t match the camera motion and results in a spatial inconsistency.

Three-view pipeline: label the sofa in V1, cut it out and inpaint V2, then paste the sofa from V3 into V2 to create an inconsistent orientation.

Models do far worse than humans and are very inconsistent across categories

All models do substantially worse than humans, and by a large margin too. This suggests that even if these models can describe spatial attributes of scenes, they lack grounded spatial understanding.

Open models
ModelAccuracy (%)
Gemma 3 12B8.5
Idefics3 8B8.9
Idefics2 8B11.9
Qwen2.5-VL 7B15.9
InternVL 3.5 8B16.1
LLaVA OneVision 1.5 8B16.6
SpaceQwen2.5-VL 3B18.2
Llama 3.2 Multimodal 11B23.4
Qwen3-VL 4B24.7
Qwen3-VL 8B Thinking25.2
Qwen3-VL 8B Instruct27.6
Ensemble (open source)30.1
Proprietary models and baselines
ModelAccuracy (%)
Random chance7.9
Human84.8
GPT-5 Nano15.3
Gemini 2.5 Flash17.6
GPT-4o19.0
Gemini 2.5 Pro (HR)28.9
Gemini 2.5 Pro (LR)29.4
Gemini 2.5 Pro (MR)29.4
GPT-5 (HR)30.2
GPT-5 (MR)31.4
GPT-5 (LR)34.2
Ensemble (all models)35.0

Overall accuracy on spatial inconsistency identification. LR, MR, and HR denote low, medium, and high reasoning budgets.

In addition, model accuracy varies greatly across object types (Qwen gets 66.7% of lamps, but 0% of floormats), while humans are very consistent. This suggests these models have a very brittle understanding of spatial consistency.

Accuracy across object categories: humans outperform the three highlighted models in every category, while model performance varies widely. Dashed lines show random chance.

Models don’t cheat

We empirically find that artifacts from cut-and-paste do not confound our results.

Accuracy (%)
ModelBenchmarkSelf-pasteNo change
GPT-534.215.113.7
Gemini 2.5 Pro29.413.815.3
Qwen3-VL 8B27.621.322.0

Notably, models select a cut-and-pasted object with a different pose (benchmark) far more than if we cut-and-paste the same pose. This suggests models look for 3D structure and not artifacts.

More reasoning doesn’t help

Accuracy (%) by reasoning budget
ModelLowMediumHigh
GPT-534.231.430.2
Gemini 2.5 Pro29.429.428.9

Increasing the reasoning budget does not improve spatial reasoning. This is counterintuitive and suggests that models’ reasoning abilities do not represent spatial information well.