BoundInk treats inter-character boundaries as explicit generation units, cutting normalized DTW by 17.6–47.8% and winning 78.0–82.6% of blind criterion-wise judgments
Synopsis
BoundInk is a writer-conditioned online handwriting generation framework that treats inter-character boundaries (cursive joins, spacing, alignment) as explicit generation targets, modeling predecessor-to-current local transitions with a bigram-aware sliding-window Transformer while injecting sentence context through gated fusion, and it introduces Connectivity and Spacing Metrics (CSM) to directly assess cursive continuity and character/word spacing; across three benchmark-matched settings it improves all applicable boundary-quality measures, reduces normalized DTW by 17.6–47.8%, and is preferred in 78.0–82.6% of valid criterion-wise judgments in blind human evaluation.
Interpretation
It elevates inter-character boundaries from incidental artifacts of long-sequence decoding to explicit generation units: BoundInk uses a bigram-aware sliding-window Transformer decoder with causal self-attention and RoPE over a predecessor–current window, while sentence-level information enters through position-specific Context Memory and token-wise gated fusion, so writer-specific glyph appearance is preserved while connections, spacing, and alignment adapt. Prior methods such as DeepWriting, DSD, and OLHWG capture inter-character behavior only implicitly through long-range sequence modeling, and OLHWG models layout and spacing but not pen-down continuation across characters; BoundInk separates local boundary dynamics from sentence context into two complementary pathways. Ablation shows that removing Bi-SWT drops CSM F1 from 0.49 to 0.37, SSS from 0.71 to 0.40, and raises normalized DTW from 0.51 to 0.60; removing gated context fusion mainly affects cursive connectivity (F1 from 0.49 to 0.35); removing both yields F1 of 0.22.
It introduces Connectivity and Spacing Metrics (CSM), which directly measure cursive connectivity, inter-character spacing, and word spacing as a complement to global trajectory similarity. Conventional trajectory distances such as DTW can score a globally similar output well even when local joins are wrong or spacing is unnatural; CSM comprises F1, CRE, KGS, and SSS, aggregated per writer and then macro-averaged, with KGS and SSS using an overlap-aware symmetric log-ratio similarity. The paper presents qualitative cases where DTW differences are small while CSM differences are pronounced, and explicitly notes these are selected examples rather than evidence that CSM is generally anti-correlated with DTW.
Across three benchmark-matched protocols, BoundInk improves all applicable boundary-quality measures and reduces normalized DTW. Against DeepWriting normalized DTW falls from 2.70 to 1.41 and against DSD from 1.25 to 0.99; under the OLHWG-compatible protocol KGS rises from 0.34 to 0.40 and normalized DTW falls from 0.17 to 0.14; glyph-level DTW drops from 1.6048 to 1.3649 in English and from 0.8789 to 0.8718 in Chinese. BoundInk is retrained for each baseline with the corresponding split, preprocessing, and metric implementation, and the comparison set is compatibility-based rather than exhaustive; CSM components are reported as undefined when a protocol lacks the required boundary type.
Blind human preference and content/writer verification support that boundary-quality gains do not come at the expense of legibility or writer consistency. Thirty participants each completed 20 randomized trials, producing 1,800 criterion-level decisions of which 1,604 were valid and 196 were marked Cannot judge; BoundInk received 615 of 789 valid preferences against DSD and 673 of 815 against DeepWriting. OCR-based recognition on rendered trajectories remains competitive with baselines, while writer-classification Top-1 rises from 35.8 to 61.7 against DeepWriting and from 6.0 to 69.0 against DSD, with Top-5 rising from 68.3 to 81.7 and from 27.0 to 74.0 respectively.
Perspective
The work targets writer-conditioned sentence-level online handwriting generation in settings with character-boundary annotations and writer references, covering English (IAM-expanded, BRUSH, and a merged IAM–BRUSH set) and Chinese (CASIA-OLHWDB 2.0–2.2); style conditioning uses image-based character references, so no online trajectory sequence from the writer is needed at inference. For downstream use, CSM offers a directly measurable tool for cursive connectivity, inter-character spacing, and word spacing when comparing generators' boundary behavior; the authors also list optional sequence references, broader scripts, unseen character combinations, scarce or mismatched references, long-sequence decoding efficiency, and autoregressive error accumulation as next directions.
The merged English dataset is used for qualitative examples, human evaluation, and ablations, while benchmark-matched quantitative comparisons retrain BoundInk under each baseline protocol, so numbers across protocols should not be compared directly. CSM components are reported as undefined when a protocol lacks the required boundary type, for example only KGS is defined under the OLHWG-compatible protocol. The V2 exploratory analysis in the appendix compares two model revisions at an intermediate checkpoint, and the authors state it is not a controlled single-factor ablation and that the strictly aligned subset covers only part of the generated test sentences. The authors also note that standardized benchmarks are still needed to jointly assess content, style, boundary quality, reliability, and efficiency.
