Skip to content
GeoSpatialBench
BenchmarkGeospatial reasoning · Spatial editing

Can models build the spatial relations they recognize?

Spatial intelligence requires more than locating objects in space.

AI systems must also reason about geometric constraints and construct spatial configurations that satisfy them.

GeoSpatialBench evaluates this capability across 28 classical spatial relations and six real-world cities, using both spatial QA and constrained geometric editing.

Editing task · Metric
Move streetlight F0 so its Euclidean distance to zone F1 is 8–22 m, while remaining disjoint.

F1 · zone VALID BAND · 8–22 M FROM F1 F0 BEFORE 15.0 m F0 N 20 M
KIMI K3, same instruction, two inputs. Outcomes from the paper’s Euclidean-distance editing case; geometry is schematic.
8,400instances
28relations in 4 families
6cities
3input modalities
9models evaluated
Motivation

Why Spatial Editing?

From Spatial Recognition to Geometric Action

A model may correctly identify whether two roads are connected, which building lies to the north, or how far apart two objects are. But can it modify a spatial configuration to make a required relation hold?

This distinction matters for AI systems expected to plan, construct, and modify environments, where geometric constraints must be satisfied by the resulting configuration.

Spatial QA measures important aspects of spatial understanding, but provides limited evidence of this constructive capability. GeoSpatialBench extends spatial evaluation from recognizing relations to constructing geometrically valid outcomes.

Recognize

“Do F0 and F1 touch?”

Model reads the geometry and names the relation.
92.7%MCQ-4
Construct

“Move F0 so it touches F1.”

Model must output geometry that makes the relation hold.
28.8%Editing CSR
Touching (T04), GeoJSON input, pooled over nine models (Table 4). Prompts are illustrative.
Key findings

What the benchmark shows

Five key findings across nine models. Explore the supporting results in each chart.

Finding Evidence / key result Why it matters
01

Explicit geometry supports precise reasoning better than map images

GPT-6 Astra · QA: GeoJSON 96.5% → Map image 66.8%

Every model scores higher with GeoJSON; the map costs 18.1% in QA on average.

Coarse judgments survive on a rendered map, but exact boundaries and distances do not. Maps remain an unreliable medium for precise spatial reasoning.

02

Spatial representation shapes reasoning reliability

Editing · GeoJSON + Map vs. GeoJSON: −10.1 to +8.4%

QA changes stay within approximately ±3%. Adding the map brings no consistent gain.

Reliable spatial AI may benefit from representations that explicitly preserve geometric structure, beyond visual information alone.

03

Recognizing a relation does not guarantee constructing it

Touching · GeoJSON: MCQ-4 92.7% → Editing 28.8%

MiMo-V2.6-Pro answers better than Qwen3.8-27B but edits worse.

High spatial QA accuracy does not guarantee successful geometric action. Editing reveals capabilities that recognition-only evaluation can miss.

04

Geometric precision remains a bottleneck

Metric · GeoJSON: Open QA 38.6% · MCQ-4 53.4% · Editing 42.9%

Metric relations are the lowest family in every task and input.

Providing geometry does not ensure accurate geometric computation. Precise quantitative reasoning remains an important direction for spatial AI.

05

Scores vary far more between models than between cities

Between models: 67.0% · Within a model: 1.0–5.7%

Open-ended GeoJSON QA across the six cities.

The differences reflect the models rather than the city tested. Six cities do not, however, establish generalization to unseen ones.

Benchmark

One set of relations, tested two ways

A correct answer does not show that a model can apply the relation. GeoSpatialBench poses the same 28 relations as questions and as edits, so both abilities are measured on the same set of relations. QA and editing draw on separate instance pools.

Source

OpenStreetMap · Six Cities

Stockholm · Paris · Nairobi · Hong Kong · Wuhan · Detroit
Roads, buildings, land use, and points of interest, projected to a city-specific metric coordinate system and filtered by type, length, and area.

Scene

Canonical scene + geometric oracle

Each instance records the geometries, roles and attributes it needs. An oracle computes the reference answer. Source identities are kept for auditing and hidden from models.

Inputs

The same scene, three ways

  • GeoJSON: exact coordinates, anonymous feature IDs
  • Map image: annotated Web Mercator map, 1,536 × 1,536 px
  • Together: the map plus the same GeoJSON
Tasks

Answer, then edit

  • QA: 4,200 questions, open-ended and four-option multiple choice. Answers are a category, a number, a feature ID or an ordering.
  • Editing: 4,200 instructions. Change only the target feature F0 so that every constraint holds. Instructions whose constraints already hold are discarded, and a reference edit confirms that each is solvable.

28 relations, four families

Each relation is defined by a precise geometric or attribute rule drawn from GIS and spatial-database standards. Two roads count as connected only if they share an endpoint.

Scoring QA

Scoring is deterministic, with no learned judge. Distance and perimeter answers have a 1 m tolerance, bearings 1° (circular) and density 0.001. Unparseable answers count as wrong.

Open-ended and multiple-choice accuracy are reported separately.

Scoring structured edits

GeoJSON and Together models return a transformation (translate, rotate, scale or set an attribute) and an executor applies it to F0. An edit succeeds only if it is valid, actually changes F0, and satisfies every constraint. Any such edit counts, not only the reference.

One corrective retry is allowed for undecodable or unchanged edits; it does not reveal whether the constraints were met.

Scoring map-image edits

Image models return edited geometry in pixel coordinates, which is mapped back to the task coordinates. An edit passes if it satisfies the constraints in metric space or falls within 4 px Hausdorff distance of the reference.

This raster-tolerant criterion changes both input and output, so the gap from GeoJSON is not a pure input ablation.

Constraint Satisfaction Rate

Structured edits · GeoJSON and Together inputs

Vi
1 if the prediction is parseable, nonempty and valid
Ui
1 if a geometric edit changes F0 (always 1 for attribute edits)
Ci
the constraints of instance i
Si, S ′i
the scene before and after the edit
700 QA + 700 editing instances per city StockholmParisNairobiHong KongWuhanDetroit
Results

Nine models, three inputs

Five API models and four open-weight models. Every score covers 4,200 instances across the six cities.

01Input representation

Explicit geometry supports precise reasoning better than map images.

Every model scores higher from GeoJSON than from the map image. The map costs most where exact values are needed: some distance and density questions are never answered exactly from it.

−18.1%

QA, map image vs. GeoJSON
mean of nine models; every model is lower

38.6 → 15.1%

Metric relations, open QA
GeoJSON → map image

0.0%

Minimum distance, Manhattan distance and POI density, open QA from the map
vs. 27.8%, 44.7% and 10.2% from GeoJSON

GeoJSONMap imageTogether (map + GeoJSON)
One row per model, sorted by GeoJSON open-ended QA; hover for values. Together is discussed in Finding 02. Map-image editing returns pixel coordinates and is scored by raster-tolerant success, so its editing gap is not a pure input effect. Source: Tables 1 and 4.
Table view Table 1 · accuracy / success, %
02Combined input

Spatial representation shapes reasoning reliability.

Adding the map to GeoJSON does not consistently improve performance. QA barely moves. In editing, three models gain and six lose.

−0.2%

Mean QA change with the map added
Together − GeoJSON; −2.1 to +3.1 per model

−2.2%

Mean editing change with the map added
Together − GeoJSON, strict CSR

3 of 9

models edit better with the map added
largest gain: MiMo-V2.6-Pro, +8.4

Change = Together − GeoJSON, in percentage points. QA is the mean of open-ended and MCQ-4 accuracy; both editing conditions use strict CSR with the same structured interface. Differences are between pooled scores; no significance tests are reported. Source: Table 1.
03Recognition vs. editing

Recognizing a relation does not guarantee constructing it.

Relations that models recognize almost perfectly can still be hard to construct. How large the gap is depends on the relation.

T04 Touching92.7 → 28.8

MCQ-4 → editing, GeoJSON

T02 Disjointness91.6 → 52.7

MCQ-4 → editing, GeoJSON

T01 Equality98.4 → 96.7

MCQ-4 → editing, GeoJSON

Across modelsGeoJSON · QA mean vs. editing CSR

Mostly aligned, with exceptions. MiMo-V2.6-Pro answers better than Qwen3.8-27B (67.3% vs. 55.5%) but edits worse (43.0% vs. 54.4%).

Across relationsGeoJSON · pooled over nine models
Diagonal: equal QA and editing scores. Points in the shaded half are recognized better than they are constructed; each stem shows the gap. Relation scores are 100 − error rate (Table 4).
04Relation families

Geometric precision remains a bottleneck.

Metric relations score lowest in every task and input. Exact coordinates do not guarantee exact computation.

38.6%

Metric relations, open QA
GeoJSON · lowest family

42.9%

Metric relations, editing
GeoJSON · lowest family

10.2%

POI density, open QA
GeoJSON · lowest of all 28 relations

GeoJSONMap image
Accuracy / success, % · pooled over nine models
Accuracy = 100 − error rate (Table 3). Map-image editing uses the raster-tolerant criterion.
All 28 relations Table 4 as accuracy · click a column to sort
05Geography

Scores vary far more between models than between cities.

Every model performs about the same in all six cities. The differences come from the models, not the test locations.

67.03%

between the strongest and weakest model
open QA, GeoJSON

1.00–5.71%

spread across six cities, within one model
open QA, GeoJSON

Dots: the six cities (700 questions each). Bar: pooled accuracy (4,200). This covers the six evaluated cities only, not unseen ones. Source: Table 2.
Table view Table 2 · open-ended GeoJSON QA by city, %
Failure cases

Where it breaks

Coarse judgments survive the map; precise boundaries and distances do not. In editing, a plausible operation can still produce the wrong geometry. One QA and one editing case per family.

Question answering
Spatial editing
Conclusion

Toward Vector-Grounded Spatial Intelligence

Spatial intelligence demands more than recognizing spatial relations. It requires the ability to manipulate geometry and satisfy spatial constraints.

Our findings highlight the value of explicit vector geometry. Compared with rendered maps, GeoJSON enables more reliable spatial reasoning by exposing precise coordinates and geometric structures. It also supports executable spatial edits and deterministic constraint verification. Yet even with vector inputs, many models struggle with precise geometric reasoning and constraint satisfaction.

The next challenge is to move from reading spatial relations to acting on explicit geometry—turning spatial understanding into verifiably correct spatial configurations.

Limitations

  1. Six cities

    City-to-city variation is small for every model, but six cities do not establish generalization to unseen ones.

  2. Image editing is a different protocol

    Pixel-coordinate output and a raster-tolerant criterion mean image editing is not directly comparable to strict CSR.

  3. Attribute information

    Geometry-only inputs omit road-layer attributes. An information-sufficient score removes the 25 vertical-order questions per city that depend on them.

Resources

  • Paper

    Mind the Spatial Gap: A Benchmark for Geospatial Reasoning and Constrained Spatial Editing

    PDF
  • Evaluation code

    Answer parsing, numerical tolerances, the edit executor and GIS constraint checks.

    On request
  • Benchmark data

    Canonical scenes as GeoJSON and annotated maps, prompts, reference answers and metadata for all 8,400 instances.

    On request