Can models build the spatial relations they recognize?
Spatial intelligence requires more than locating objects in space.
AI systems must also reason about geometric constraints and construct spatial configurations that satisfy them.
GeoSpatialBench evaluates this capability across 28 classical spatial relations and six real-world cities, using both spatial QA and constrained geometric editing.
Editing task · Metric
Move streetlight F0 so its Euclidean distance to zone F1 is 8–22 m, while remaining disjoint.
Why Spatial Editing?
From Spatial Recognition to Geometric Action
A model may correctly identify whether two roads are connected, which building lies to the north, or how far apart two objects are. But can it modify a spatial configuration to make a required relation hold?
This distinction matters for AI systems expected to plan, construct, and modify environments, where geometric constraints must be satisfied by the resulting configuration.
Spatial QA measures important aspects of spatial understanding, but provides limited evidence of this constructive capability. GeoSpatialBench extends spatial evaluation from recognizing relations to constructing geometrically valid outcomes.
“Do F0 and F1 touch?”
Model reads the geometry and names the relation.“Move F0 so it touches F1.”
Model must output geometry that makes the relation hold.What the benchmark shows
Five key findings across nine models. Explore the supporting results in each chart.
Explicit geometry supports precise reasoning better than map images
Every model scores higher with GeoJSON; the map costs 18.1% in QA on average.
Coarse judgments survive on a rendered map, but exact boundaries and distances do not. Maps remain an unreliable medium for precise spatial reasoning.
Spatial representation shapes reasoning reliability
QA changes stay within approximately ±3%. Adding the map brings no consistent gain.
Reliable spatial AI may benefit from representations that explicitly preserve geometric structure, beyond visual information alone.
Recognizing a relation does not guarantee constructing it
MiMo-V2.6-Pro answers better than Qwen3.8-27B but edits worse.
High spatial QA accuracy does not guarantee successful geometric action. Editing reveals capabilities that recognition-only evaluation can miss.
Geometric precision remains a bottleneck
Metric relations are the lowest family in every task and input.
Providing geometry does not ensure accurate geometric computation. Precise quantitative reasoning remains an important direction for spatial AI.
Scores vary far more between models than between cities
Open-ended GeoJSON QA across the six cities.
The differences reflect the models rather than the city tested. Six cities do not, however, establish generalization to unseen ones.
One set of relations, tested two ways
A correct answer does not show that a model can apply the relation. GeoSpatialBench poses the same 28 relations as questions and as edits, so both abilities are measured on the same set of relations. QA and editing draw on separate instance pools.
OpenStreetMap · Six Cities
Stockholm · Paris · Nairobi · Hong Kong · Wuhan · Detroit
Roads, buildings, land use, and points of interest,
projected to a city-specific metric coordinate system
and filtered by type, length, and area.
Canonical scene + geometric oracle
Each instance records the geometries, roles and attributes it needs. An oracle computes the reference answer. Source identities are kept for auditing and hidden from models.
The same scene, three ways
- GeoJSON: exact coordinates, anonymous feature IDs
- Map image: annotated Web Mercator map, 1,536 × 1,536 px
- Together: the map plus the same GeoJSON
Answer, then edit
- QA: 4,200 questions, open-ended and four-option multiple choice. Answers are a category, a number, a feature ID or an ordering.
- Editing: 4,200 instructions. Change only the target feature F0 so that every constraint holds. Instructions whose constraints already hold are discarded, and a reference edit confirms that each is solvable.
28 relations, four families
Each relation is defined by a precise geometric or attribute rule drawn from GIS and spatial-database standards. Two roads count as connected only if they share an endpoint.
Scoring QA
Scoring is deterministic, with no learned judge. Distance and perimeter answers have a 1 m tolerance, bearings 1° (circular) and density 0.001. Unparseable answers count as wrong.
Open-ended and multiple-choice accuracy are reported separately.
Scoring structured edits
GeoJSON and Together models return a transformation (translate, rotate, scale or set an attribute) and an executor applies it to F0. An edit succeeds only if it is valid, actually changes F0, and satisfies every constraint. Any such edit counts, not only the reference.
One corrective retry is allowed for undecodable or unchanged edits; it does not reveal whether the constraints were met.
Scoring map-image edits
Image models return edited geometry in pixel coordinates, which is mapped back to the task coordinates. An edit passes if it satisfies the constraints in metric space or falls within 4 px Hausdorff distance of the reference.
This raster-tolerant criterion changes both input and output, so the gap from GeoJSON is not a pure input ablation.
Structured edits · GeoJSON and Together inputs
- Vi
- 1 if the prediction is parseable, nonempty and valid
- Ui
- 1 if a geometric edit changes F0 (always 1 for attribute edits)
- Ci
- the constraints of instance i
- Si, S ′i
- the scene before and after the edit
Nine models, three inputs
Five API models and four open-weight models. Every score covers 4,200 instances across the six cities.
Explicit geometry supports precise reasoning better than map images.
Every model scores higher from GeoJSON than from the map image. The map costs most where exact values are needed: some distance and density questions are never answered exactly from it.
QA, map image vs. GeoJSON
mean of nine models; every model is lower
Metric relations, open QA
GeoJSON → map image
Minimum distance, Manhattan distance and POI density, open QA from the map
vs. 27.8%, 44.7% and 10.2% from GeoJSON
Table view Table 1 · accuracy / success, %
Spatial representation shapes reasoning reliability.
Adding the map to GeoJSON does not consistently improve performance. QA barely moves. In editing, three models gain and six lose.
Mean QA change with the map added
Together − GeoJSON; −2.1 to +3.1 per model
Mean editing change with the map added
Together − GeoJSON, strict CSR
models edit better with the map added
largest gain: MiMo-V2.6-Pro, +8.4
Recognizing a relation does not guarantee constructing it.
Relations that models recognize almost perfectly can still be hard to construct. How large the gap is depends on the relation.
MCQ-4 → editing, GeoJSON
MCQ-4 → editing, GeoJSON
MCQ-4 → editing, GeoJSON
Mostly aligned, with exceptions. MiMo-V2.6-Pro answers better than Qwen3.8-27B (67.3% vs. 55.5%) but edits worse (43.0% vs. 54.4%).
Geometric precision remains a bottleneck.
Metric relations score lowest in every task and input. Exact coordinates do not guarantee exact computation.
Metric relations, open QA
GeoJSON · lowest family
Metric relations, editing
GeoJSON · lowest family
POI density, open QA
GeoJSON · lowest of all 28 relations
All 28 relations Table 4 as accuracy · click a column to sort
Scores vary far more between models than between cities.
Every model performs about the same in all six cities. The differences come from the models, not the test locations.
between the strongest and weakest model
open QA, GeoJSON
spread across six cities, within one model
open QA, GeoJSON
Table view Table 2 · open-ended GeoJSON QA by city, %
Where it breaks
Coarse judgments survive the map; precise boundaries and distances do not. In editing, a plausible operation can still produce the wrong geometry. One QA and one editing case per family.
Toward Vector-Grounded Spatial Intelligence
Spatial intelligence demands more than recognizing spatial relations. It requires the ability to manipulate geometry and satisfy spatial constraints.
Our findings highlight the value of explicit vector geometry. Compared with rendered maps, GeoJSON enables more reliable spatial reasoning by exposing precise coordinates and geometric structures. It also supports executable spatial edits and deterministic constraint verification. Yet even with vector inputs, many models struggle with precise geometric reasoning and constraint satisfaction.
The next challenge is to move from reading spatial relations to acting on explicit geometry—turning spatial understanding into verifiably correct spatial configurations.
Limitations
- Six cities
City-to-city variation is small for every model, but six cities do not establish generalization to unseen ones.
- Image editing is a different protocol
Pixel-coordinate output and a raster-tolerant criterion mean image editing is not directly comparable to strict CSR.
- Attribute information
Geometry-only inputs omit road-layer attributes. An information-sufficient score removes the 25 vertical-order questions per city that depend on them.
Resources
-
Paper
Mind the Spatial Gap: A Benchmark for Geospatial Reasoning and Constrained Spatial Editing
PDF - Evaluation code
Answer parsing, numerical tolerances, the edit executor and GIS constraint checks.
On request - Benchmark data
Canonical scenes as GeoJSON and annotated maps, prompts, reference answers and metadata for all 8,400 instances.
On request