Skip to main content
Back to timeline
arXivSource publication:

AbSteering steers general-purpose VideoLMs with abnormality-centric chain-of-thought and DPO to generate HRCT reports, surpassing large-scale CT-specific foundation models on fine-grained clinical metrics while improving detection sensitivity and reducing hallucinations.

Synopsis

The work presents AbSteering, a two-stage framework combining abnormality-centric chain-of-thought training with a Direct Preference Optimization objective for fine-grained abnormality discrimination, to adapt general-purpose VideoLMs to high-resolution CT report generation, and curates the CT-RATE-AB dataset; results show that general-purpose VideoLMs transfer effectively to 3D medical imaging under limited data, achieving state-of-the-art performance on fine-grained clinical efficacy metrics, with superior detection sensitivity over domain-specific CT foundation models pretrained on large-scale CTs while mitigating hallucinations.

Source-provided article image: Unleashing Video Language Models for Fine-Grained HRCT Report Generation

Interpretation

General-purpose VideoLMs transfer effectively to high-volume 3D medical imaging report generation under limited data, offering an efficient alternative to training modality-specific foundation models from scratch. Prior CT report generation largely relied on training or heavily tailoring modality-specific encoders, which is data- and compute-intensive; this work systematically studies cross-modal transferability of general-purpose VideoLMs by viewing an HRCT volume as a video-like slice sequence. The paper lists cross-modal transferability as one of three contributions and reports in the abstract that general-purpose VideoLMs possess strong transferability when guided by this paradigm; specific data scale and control settings require the original experiments.

Abnormality-centric chain-of-thought training reformulates the vision-to-text task into a structured sequence transitioning from abnormality detection to report composition, compelling explicit reasoning over pathological findings and learning standardized report structures. Unlike direct report generation, the method first normalizes CT-RATE reports into a unified (region: abnormality) template covering ten anatomical regions, with two auxiliary categories, uncategorized and repetitive, introduced for verification. The method section specifies the ten-region anatomical taxonomy and a GPT-4o sentence-level assignment process, manually refined into the CT-RATE-AB dataset; quantitative effects require the original experiments.

Fine-grained abnormality discrimination via Direct Preference Optimization uses clinically confusable abnormalities within the same anatomical region as hard negatives, enhancing discrimination of subtle pathological differences and reducing hallucinations. The work introduces preference optimization into CT report generation, constructing hard negatives from clinically confusable abnormalities to target long-tail, localized abnormality recognition where existing approaches remain limited. The abstract reports that AbSteering achieves state-of-the-art performance on fine-grained clinical efficacy metrics, with superior detection sensitivity over domain-specific foundation models pretrained on large-scale CTs while mitigating hallucinations; specific metric values require the original text.

The CT-RATE-AB dataset is curated to facilitate chain-of-thought training in medical imaging and to evaluate diversity and fine-grained abnormality recognition for high-volume chest CT understanding. Built on CT-RATE, the dataset reorganizes reports according to the anatomical taxonomy and provides structured-to-raw report pairings, serving chain-of-thought training and fine-grained evaluation. The paper lists the dataset as one of three contributions and states that data and model weights are released; dataset scale and evaluation protocol details require the original text.

Perspective

The work targets high-volume chest HRCT report generation and is suited to researchers and engineering teams seeking to adapt general-purpose VideoLMs to 3D medical image interpretation with limited data; its abnormality-centric chain-of-thought relies on report templates restructured by a ten-region anatomical taxonomy, and its preference optimization relies on hard negatives formed by clinically confusable abnormalities within the same anatomical region, so the method presupposes structurable report data and definable confusable abnormality pairs. The released data and model weights provide a starting point for subsequent transfer studies on other volumetric modalities or anatomical regions.

The loaded text is the paper body and does not include full experimental tables or specific metric values, so the exact numbers for detection sensitivity, hallucination reduction, and comparison baselines cannot be verified here; the scale of CT-RATE-AB, annotation consistency, and the separate ablation contributions of chain-of-thought and preference optimization still require the original experiments. In addition, how general-purpose VideoLMs transfer to other volumetric modalities, other anatomical regions, and different reporting conventions remains an open question for follow-up work.

Sources