Skip to main content
Back to timeline
GIScience & Remote SensingSource publication:

After step-by-step exploration of 20×20 symbolic maps, node-sequence memory lifts GPT-5.2 total accuracy from 43.89% to 77.78%, while further model versions and parameter scale add little spatial reasoning

Synopsis

The study proposes an interactive evaluation framework in which foundation model agents incrementally explore partially observable 20×20 grid-based symbolic maps of roads, intersections, and POIs, then probes spatial understanding with direction judgment, distance estimation, proximity judgment, POI density recognition, and path planning; by systematically varying exploration strategies, memory representations, and reasoning prompts, it finds that exploration has limited impact on final reasoning accuracy, that memory representation (especially node-sequence and graph memory) is central, that structured memory and advanced prompts repair reasoning failures through explicit spatial reconstruction, and that spatial reasoning performance saturates across model versions and scales beyond a cap

Source-provided article image: Thinking on maps: how foundation model agents explore, remember, and reason in map environments
Figure 1

Figure 1. Framework

· Page 9

Interpretation

It proposes an interactive, experience-driven evaluation framework in which agents explore symbolic maps step by step and build internal representations from local observations rather than receiving a complete map. Prior evaluations of foundation model spatial ability largely relied on static map inputs or text-based queries, treating the model as an external reasoner; this framework embeds the model as an incremental explorer. The framework is built on OpenStreetMap data from 15 cities across China, the USA, and Europe, with all maps normalized to 20×20 grids retaining 9 to 21 POIs each, averaging 15.27 POIs, 20.93 road intersections, and 7.67 main roads.

Memory representation is the central factor in spatial reasoning: Node-Sequence Memory (NSM) achieves the highest total accuracy for all five models, ranging from 55.83% to 86.11%, versus 43.61% to 79.44% under Simple Dialogue Memory (SDM). Compared with the bounded differences produced by exploration strategies, changing memory structure induces substantially larger performance variation, indicating that how experience is organized matters more than how it is sampled. With Nearest-POI Strategy and simple prompting fixed, GPT-5.2's distance estimation rises from 20.00% under SDM to 81.67% under NSM and path planning from 50.00% to 91.67%; Claude-4.5's path planning rises from 58.33% to 95.00%.

Different memory structures show task-dependent trade-offs: Graph Memory and Map Memory are strong on path planning (91.67% to 95.00% for GPT-5.2, Gemini-2.5-Pro, and Claude-4.5) but clearly weaker than NSM on direction judgment and distance estimation. This shows that more abstract memory is not uniformly better; its gains concentrate on structural tasks while attenuating fine-grained directional and metric cues. GPT-5.2 under Graph Memory reaches only 31.67% on direction judgment and 33.33% on distance estimation, versus 75.00% and 81.67% under NSM; the hybrid NSM+SDM yields the largest overall gains, lifting GPT-5.2 from 43.89% to 77.50%.

Reasoning prompts act mainly as capability amplifiers: Tree-of-Thoughts (ToT) achieves the highest total accuracy for all five models, improving on Default Thought by 3 to 18 percentage points, with larger gains for weaker-baseline models. Prompting does not alter task-intrinsic difficulty or compensate for missing spatial evidence; it improves how existing spatial knowledge is utilized and stabilizes multi-step inference. DeepSeek-V3.2 rises from 55.83% under DT to 73.89% under ToT and Qwen-3 from 62.22% to 72.22%, while Gemini-2.5-Pro stays above 84% under all prompts with little change.

Perspective

The framework targets symbolic, discretized 20×20 grid maps where roads and POIs are categorical symbols, POIs link to their nearest road cell, agents move in eight directions under a 5×5 local view, and travel between POIs always follows the shortest road path. It suits controlled comparison of spatial understanding across cities and tasks, serving readers who study map-based spatial cognition and design map agents. The study also reports bit-level memory sizes: SDM exceeds 20,000 bits yet yields the lowest accuracy, NSM sits around 11,115.3 to 11,519.3 bits with the highest accuracy, GM around 2,730.8 to 3,122.6 bits, and MM around 3,704.3 to 3,844.6 bits, suggesting an intermediate region that reduces redundancy while preserving sequential information.

The conclusions rest on symbolic, planar grid maps with shortest-path navigation, and the authors note this simplifies geometric detail, scale variation, and cartographic conventions, with future work extending to multi-scale or vector map representations. Memory structures are explicitly engineered rather than learned or adaptive, so the evaluation concerns how models use given structured memory rather than how such memory could be autonomously formed or refined. Task accuracy is the primary metric, leaving reasoning efficiency, robustness, and error patterns outside the current assessment. In addition, the saturation finding across model versions and parameter sizes comes from comparing multiple Claude versions and Qwen parameter sizes, and its scope merits continued observation across more model families.

Sources