Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Simulation-to-agent VLM framework: video memory lifts wildfire four-tag accuracy from 22.6% to 51.5%, full system reaches 77.3% on report fields

The work presents a simulation-grounded vision-language model framework that automatically converts 2D wildfire simulations into labeled video episodes via a fixed Blender mapping, and uses them as reusable multimodal memory for a training-free multi-agent VLM system that retrieves reference episodes, reconciles visual and memory-based predictions, and produces structured wildfire reports; on held-out generated episodes, video memory achieves 51.5% exact four-tag accuracy versus 22.6% for direct VLM querying and 16-17% for text-only memory, while the complete system reaches 77.3% accuracy on six simulator-derived report fields, with component ablations, cross-generator tests, and three real-UAV evaluations assessing retrieval, reporting, generator changes, and observable monitoring tasks.