Simulation-to-agent VLM framework: video memory lifts wildfire four-tag accuracy from 22.6% to 51.5%, full system reaches 77.3% on report fields
Related research and updatesSynopsis
The work presents a simulation-grounded vision-language model framework that automatically converts 2D wildfire simulations into labeled video episodes via a fixed Blender mapping, and uses them as reusable multimodal memory for a training-free multi-agent VLM system that retrieves reference episodes, reconciles visual and memory-based predictions, and produces structured wildfire reports; on held-out generated episodes, video memory achieves 51.5% exact four-tag accuracy versus 22.6% for direct VLM querying and 16-17% for text-only memory, while the complete system reaches 77.3% accuracy on six simulator-derived report fields, with component ablations, cross-generator tests, and three real-UAV evaluations assessing retrieval, reporting, generator changes, and observable monitoring tasks.
Figure 4: We show the effects of different multimodal embedding models on tag retrieval accuracy. From experiments, Qwen3-2B-VL gives the best retrieval accuracy in most cases which leads to our choice of it being our embedding model.
arXivInterpretation
An automatic pipeline converts 2D wildfire simulations into labeled video episodes: a fixed Blender mapping produces low-detail 3D proxies aligned with simulator terrain, fuel layout, fire activity, and wind cues, while controllable video generation supplies richer appearance. Relating visual evidence to physical fire dynamics previously required real videos with synchronized physical annotations, which are scarce, and high-fidelity 3D simulation is costly; this pipeline pairs simulator labels with generated videos to form reusable multimodal memory. The abstract states the proxies are intermediate representations rather than finely rendered final scenes, and reports that on held-out generated episodes video memory achieves 51.5% exact four-tag accuracy, above 22.6% for direct VLM querying and 16-17% for text-only memory.
A training-free multi-agent VLM system retrieves reference episodes, reconciles visual and memory-based predictions, and produces structured wildfire reports. Compared with direct VLM querying or text-only memory, introducing video memory and multi-agent reconciliation yields higher accuracy on simulator-derived report fields. The complete system achieves 77.3% accuracy on six simulator-derived report fields, and component ablations assess the contributions of retrieval and reporting.
Cross-generator tests and three real-UAV evaluations assess performance under generator changes and observable monitoring tasks. Evaluation extends beyond a single generator to different generators and to real-UAV scenarios for observable monitoring tasks, adding applicability evidence for the simulation-to-proxy pipeline. The abstract reports cross-generator tests and three real-UAV evaluations assessing retrieval, reporting, generator changes, and observable monitoring tasks.
Perspective
The framework targets wildfire monitoring where real-world physical annotations are scarce, and applies to monitoring tasks for which 2D simulations are available and structured reports are needed; its proxies are intermediate representations rather than finely rendered scenes, so results apply to the simulator-derived label system and report fields. For researchers and engineering teams seeking to reduce reliance on synchronized physically annotated real videos, the work offers reusable multimodal memory and a training-free multi-agent reporting pipeline, and it has been examined in three real-UAV evaluations of observable monitoring tasks.
The abstract does not state the size of the held-out generated episodes, the exact definitions of the four tags and six fields, the number of generators used in cross-generator tests, or the task setup and evaluation method of the three real-UAV evaluations; these affect how far the numbers 51.5%, 22.6%, 16-17%, and 77.3% can be extrapolated. In addition, because the proxies are intermediate representations rather than finely rendered scenes, their correspondence to real fire behavior still needs observation in more real settings.
