DataMagic turns raw tabular data into data videos via declarative multi-agent orchestration, lifting quality from 2.13/5 to 3.89/5 with execution success above 95%
Synopsis
DataMagic introduces the declarative specification DVSpec and a "Generate-then-Orchestrate" multi-agent strategy that authors multi-scene data videos with dynamic charts, voice narration, and synchronized animation from raw tabular data, raising video quality from 2.13/5 for the strongest direct LLM generation to 3.89/5 and execution success from 48.62%-86.24% to above 95% on 109 real-world samples, while a 12-participant user study shows about a 79.7% reduction in task time.
Interpretation
DVSpec decomposes a data video into an ordered scene sequence, each scene containing content, narration, and animation, and binds visual and animation elements through data-driven semantic references (e.g., {"sale_date": "2025-01-30"}) while replacing absolute timestamps with narration-indexed triggering. Existing declarative visualization specifications (Vega-Lite, Canis, and others) handle static charts or single-chart animations but do not cover the multi-modal, multi-scene requirements of data videos; existing data-video tools rely on hand-specified DOM identifiers and absolute timestamps that break when data or narration changes. The paper formalizes a scene 5-tuple, an animation 4-tuple, and narration-indexed triggering, and implements them in the system; the ablation shows that even without global orchestration the Animation dimension stays at 3.95, which the authors use to argue synchronization is guaranteed at the declarative level by DVSpec.
The "Generate-then-Orchestrate" strategy first has Story Planner, Data Manager, and Visual Designer produce a candidate scene pool in parallel, then has the Narration Director and Animation Coordinator perform global scene selection, ordering, context-aware narration generation, and audio-visual binding. Prior AI-assisted tools mostly generate individual components (charts or scripts) with limited support for organizing multiple scenes into globally coherent narratives, and greedy scene-by-scene generation tends to produce redundancy or narrative discontinuities. Ablation with Claude-Sonnet-4 as the base model: removing the Story Planner drops the average score from 3.89 to 3.44 (-11.6%), and removing Orchestration drops it to 3.54 (-9.0%), with Intent falling 12.4% and Narrative 10.9%.
The system exposes DVSpec as a shared editable state supporting three interaction modes - canvas manipulation, structured script editing, and natural language commands - where edits are scoped to a target scene and re-rendered incrementally without rerunning the full pipeline. Pixel-level generation models treat a video as a continuous frame stream with opaque intermediate results, so local edits require regenerating the whole video; DataMagic makes the scope of each modification explicit and controllable. In the DataMagic condition each of the 12 participants completed at least one intended edit, and all 12 successfully applied their requested changes within the target scene without disturbing the rest of the video.
On 109 real-world samples, DataMagic raises average quality from 1.91-2.22 for direct generation baselines to 3.38-3.89 and stabilizes execution success above 95%; in the user study, task time falls from 39.2 minutes to 8.0 minutes. Direct generation methods show high variance in execution rates (48.62%-86.24%) and are weakest on animation and narrative; DataMagic shows its largest gains in exactly those two dimensions (Animation +120%, Narrative +72%). The evaluation set draws on T2R-bench and DAComp-DA, totaling 60 datasets and 109 samples across 6 domains and 22 sub-domains; automated scores correlate with 3 experts on 60 stratified-sampled videos at an overall Pearson of 0.91, with all dimension-level correlations above 0.75.
Perspective
The work targets multi-scene data videos generated from single-table tabular data; the current implementation centers on data-bound statistical chart families such as bar, line, area, scatter, pie, and heatmap charts, and the available range is jointly shaped by agent prompts, underlying model capabilities, and implemented rendering components, and can be broadened by extending rendering components and generation rules. DVSpec is currently bound to a single renderer (Remotion), and specification-rendering synchronization has not been generalized across backends. The system is meant for analysts, journalists, educators, and business reporters who need to turn data analysis into video, operating under default constraints of 60-120 seconds total duration and an initial scene count, and supporting post-generation within-scene refinement through canvas, script, and natural language modes.
The data-video field still lacks a commonly accepted public benchmark, so reproducible cross-system comparison remains to be established; although the evaluation set covers 6 domains and 22 sub-domains, it is still single-table input, and multi-table joins and richer relational data are not yet included. The authors also note that structured animation patterns keep animations stable and legible but come at the cost of creative diversity, leaving open how to treat guidelines as soft constraints a model may depart from with justification, or to learn animation styles from curated exemplars. In addition, the low-quality cases in the appendix show that under extreme data distributions viewport clipping, cross-modal numerical inconsistency (narration 40.6% versus label 40.9%), and semantic gaps from the query can still occur, and these boundary conditions are not easily visible in the aggregate metrics in the main text. This reading is of the full text, but figures and appendix details are conveyed mainly through prose, so the concrete visual presentation still requires consulting the original.
