Skip to main content
Back to timeline
arXivSource publication:

TextCSP combines sub-region-aware prompts with soft cascade decoding to reach 87.0% average Dice and 4.81 mm HD95 on TextBraTS

Synopsis

The work proposes TextCSP, a hierarchical text-guided brain tumor segmentation framework built on the TextBraTS baseline with three components: a text-modulated soft cascade decoder that predicts WT→TC→ET in a coarse-to-fine manner, sub-region-aware prompt tuning that uses learnable soft prompts with a LoRA-adapted BioBERT encoder to generate branch-specialized text representations, and text-semantic channel modulators that convert those representations into channel-wise refinement signals; on the TextBraTS dataset it reaches 87.0% average Dice and 4.81 mm average HD95, improving over the previous best TextBraTS by 1.7% and about 6% (0.32 mm) respectively, with consistent gains across all three sub-regions.

Source-provided article image: Hierarchical Text-Guided Brain Tumor Segmentation via Sub-region-Aware Prompts
Fig. 1

Fig. 1. Main architecture of our TextCSP, which contains three key components: (1) a text-modulated soft cascade decoder that predicts WT→TC→ET in a coarse-to-fine manner (Sec. 2.2); (2) sub-region-aware prompt tuning (Sec. 2.3), which uses learnable soft prompts with a LoRA-adapted BioBERT encoder to generate specialized text representations tailored to each sub-region; and (3) text-semantic channel modulators that convert these representations into channel-wise refinement signals, enabling the decoder to emphasize features aligned with clinically described patterns 2.4.

· Page 3

Interpretation

A text-modulated soft cascade decoder with three independent prediction branches explicitly encodes the anatomical containment hierarchy ET⊂TC⊂WT, where coarser regions generate soft spatial attention priors for finer ones. Most prior methods use a single shared output head, ignoring the containment hierarchy and often producing inconsistent predictions such as ET outside TC; this work modulates TC features with a WT attention gate via a (1+A_WT) soft residual gate, and TC in turn constrains ET. Ablation shows the cascade strategy matters: independent parallel heads give 86.0% average Dice, partial WT→TC+ET cascading 86.7%, and the full WT→TC→ET sequential cascade 87.0%; attention visualizations show A_WT concentrating within WT toward the TC boundary and A_TC narrowing further to the ET boundary, without intermediate supervision.

Sub-region-aware prompt tuning uses three sets of learnable soft prompts (K=4, d=768) to steer one shared LoRA-adapted BioBERT encoder into branch-specialized text representations for WT, TC, and ET. Instead of compressing the whole report into a single global text embedding that overlooks how WT, TC, and ET depend on different clinical cues (edema, necrosis, enhancement), it conditions a shared encoder with different lightweight steering prompts. In ablation, adding sub-region-aware prompts to the soft cascade raises Dice from 85.9% to 86.4%, and adding LoRA further to 86.6%; prompt-length comparison gives 86.2% at K=1, 87.0% at K=4, and a slight drop to 86.7% at K=10.

Text-semantic channel modulators follow the Squeeze-and-Excitation paradigm but replace spatial squeeze with a textual squeeze, turning branch text representations into channel-wise modulation weights, applied only to the TC and ET branches. Linguistic priors are injected into decoder feature channels so the decoder emphasizes channels aligned with clinically described patterns; no text modulator is applied to WT, reflecting the design judgment that tumor-versus-background separation is primarily visual. In ablation, adding text modulation on top of the soft cascade, prompts, and LoRA raises Dice from 86.6% to 87.0% and lowers HD95 from 5.04 mm to 4.81 mm.

On the TextBraTS benchmark the method achieves consistently better results than existing methods, with 87.0% average Dice and 4.81 mm average HD95. Compared with the previous best TextBraTS (85.3% average Dice, 5.13 mm average HD95), this is a 1.7% and about 6% (0.32 mm) improvement, with gains in all three sub-regions and the largest gain for TC (+2.6%). The comparison table covers 3D-UNet, nnU-Net, SegResNet, Swin UNETR, Nestedformer, TextBraTS, and a same-platform reproduction TextBraTS†; experiments run on a single NVIDIA RTX A6000 with PyTorch and MONAI, inputs resized to 128×128, 200 training epochs with SAM and SGD plus linear warmup cosine annealing.

Perspective

The result targets brain MRI segmentation settings with paired radiological text descriptions and applies to the official train and test splits of the TextBraTS benchmark; for researchers and practitioners who want to use clinical report text to improve multi-sub-region segmentation and who care about parameter-efficient tuning (LoRA, soft prompts) and cascade decoding design, it offers a reusable component combination and ablation evidence.

Results are based on the single TextBraTS benchmark, so behavior under other data distributions, different text quality, or missing text remains unclear; the finding that prompt length K=4 beats K=1 and K=10 comes from ablations on this benchmark and its generality is still to be observed; the design judgment that text modulators are used only for TC and ET and not WT leaves open how well it holds for other tumor types or annotation conventions.

Sources