Flow-of-Thought generates intermediate visual sketches via trajectory flow fields, reaching 100% on Tetris, 99% on colored shapes, and 72.2% versus a 63.9% endpoint-only control on BLINK Multi-view
Synopsis
The work introduces Flow-of-Thought (FoT), a framework that trains coordinate-, time-, and action-conditioned spatial U-Nets with flow matching on rotation orbits and cumulative shortest-path prefixes, then freezes the learned dynamics; at inference, the source and its horizontal reflection serve as competing hypotheses selected by foreground-weighted reconstruction energy. On locked tests FoT reaches 100.0% on Tetris and 99.0% on colored shapes, with maze endpoint IoU 97.5% and 100% goal reach; under frozen transfer to 133 BLINK Multi-view public validation pairs, the orbit-trained flow reaches 72.2% versus 63.9% for its endpoint-only control.
Figure 1 : Training and inference in FoT. Top: the rotation example follows the rendered path x t = γ a ( t ) x_{t}=\gamma_{a}(t) with target transport u ∗ ( p , ω ) = ( − ω p y , ω p x ) u^{*}(p;\omega)=(-\omega p_{y},\omega p_{x}) . The field predicts u ^ t = v θ ( x t , t , c , a ) \hat{u}_{t}=v_{\theta}(x_{t},t,c,a) and minimizes ℓ ( u ^ t , u t ∗ ) \ell(\hat{u}_{t},u_{t}^{*}) ; thus only θ \theta is learned. For mazes, the same slots contain the trace τ ( t ) \tau(t) , map condition m m , and additive field v θ ( τ ( t ) , t , m ) v_{\theta}(\tau(t),t,m) . Bottom: A 𝒟 A_{\mathcal{D}} maps an arbitrary compatible item to ( x 0 , c , ℋ , 𝒜 ) (x_{0},c,\mathcal{H},\mathcal{A}) . Each hypothesis/action is integrated under the same frozen field, producing the trace { x ( t ) } \{x(t)\} and endpoint T θ ^ ( x , a ) = x ( 1 ) T_{\hat{\theta}}(x,a)=x(1) . Rotation pairs use the paper’s exact SAME/DIFFERENT energies E same E_{\rm same} and E diff E_{\rm diff} ; dense tasks return x ( 1 ) x(1) . Only A 𝒟 A_{\mathcal{D}} and the readout are dataset-specific.
arXivInterpretation
FoT treats visual sketches as intermediate reasoning steps: a coordinate-, time-, and action-conditioned spatial U-Net predicts dense transport fields, and ODE integration yields an intermediate visual state at any time rather than only a final prediction. Unlike linear pixel interpolation or endpoint-only models supervised on final transforms, rotation supervision follows the rendered orbit and maze supervision follows cumulative shortest-path prefixes, so the model learns continuous trajectories rather than a single-point mapping. Ablations hold source examples, architecture family, optimizer schedule, update count, seeds, frozen reconstruction decision, and action grid fixed: Tetris rotation orbit 100.0% versus endpoint-only 94.0% versus linear 39.0%; colored shapes 99.0% versus 85.0% versus 51.0%, with endpoint IoU improving in parallel.
After freezing the dynamics, same-versus-different rotation decisions compare the foreground-weighted reconstruction energy of two generative hypotheses, the observed source and its horizontal reflection, with no pair labels, fitted threshold, or target-task adaptation. Decision evidence comes directly from endpoints generated by the same inspectable trajectories, turning generative reconstruction energy into a discriminative answer instead of relying on semantic memorization or an extra classification head. On the locked 100-pair Tetris and colored-shape tests, FoT reaches 100.0% and 99.0%, exceeding the listed ViT, DINOv3, Qwen3.5, GPT-4o, GPT-5.6, and Claude 4.6 entries.
Under frozen transfer, the orbit-trained 2D flow reaches 72.2% on the 133-pair BLINK Multi-view public validation set, above its endpoint-only control at 63.9%, which the authors treat as the primary out-of-distribution evidence. This suggests continuous visual trajectory supervision can outperform endpoint-only supervision in some out-of-distribution settings, exceeding the listed GPT-4V at 58.7% but not the published P2+Qwen3VL at 94.0%. The two intervals are [64.0, 79.1] and [55.5, 71.6], an 8.3-point gap with overlapping intervals, which the authors call suggestive rather than conclusive; the evaluation sign convention was finalized after a two-item smoke check.
On maze navigation FoT directly generates a path, obtaining 97.5% endpoint IoU, 84.2% prefix IoU, 100% goal reach, and zero obstacle violations, while sometimes activating later path segments prematurely. Unlike baselines that emit a binary trace-level response, FoT outputs inspectable cumulative path masks, letting errors be attributed to wall avoidance, endpoint completion, or the timing of path activation. Premature activation is 31.7% and mean future-path intensity is 0.164, showing that strong endpoint and prefix metrics do not imply fully causal drawing; the authors also caution that generation metrics should not be compared directly with the trace-accuracy row.
Perspective
The framework targets spatial tasks that require step-by-step geometric transformation: same-versus-different rotation judgments and maze shortest-path generation. For a reader, it offers a reusable design pattern: learn continuous visual dynamics from trajectory supervision, freeze them, use generative reconstruction energy for discrimination, and expose intermediate frames as auditable evidence. The intended setting is in-distribution 2D rotation and maze tasks; transfer evidence comes mainly from the BLINK Multi-view public validation set, which the authors regard as the primary out-of-distribution result. The authors point to future directions including physically faithful 3D dynamics, a larger 3D test set, stricter causal path generation, and integrating the FoT pipeline as an external tool for multimodal foundation models.
A careful reader would still watch several things: the BLINK sign convention was finalized after a two-item smoke check and the authors call that result exploratory; 3D Blocks has only 78 pairs, where a single reclassified item moves accuracy by 1.3 points, the Wilson interval [50.4, 71.6] spans more than twenty points, and method rankings are seed-sensitive; and the 31.7% premature activation rate shows path generation is not fully causal. In addition, both 3D Blocks and BLINK transfers involve no target-task training, so whether the behavior holds on larger and more varied 3D test sets remains an open question.
