Fact-Flow uses LLM-bootstrapped multi-label fact guidance to raise factual accuracy of MLLM medical reports on tuberculosis and ophthalmology data
Synopsis
The work introduces Fact-Flow, a framework that decouples visual fact identification from report generation: an LLM automatically extracts and merges clinical finding labels from training reports (7 labels for tuberculosis, 42 for ophthalmology), a multi-label classifier predicts findings, and the predicted labels are serialized into a prompt to guide an MLLM; on the tuberculosis chest X-ray dataset MedGemma + Fact-Flow reaches RadFact F1 0.3055 versus 0.2266 for MedGemma alone, and on the ophthalmology dataset Qwen2.5-VL + Fact-Flow is best on most NLG metrics.
Fig. 1. The overall framework of Fact-Flow
· Page 3Interpretation
It proposes Fact-Flow, which improves MLLM-based medical report generation through explicit multi-label clinical finding conditioning, replacing the end-to-end image-to-report mapping with a two-step process via an intermediate multi-label representation Y∈{0,1}^k. Unlike generating reports directly from image features, the framework first predicts clinical findings and then generates the report from them, giving the report an explicit factual basis. Consistent improvements are reported on two datasets and three MLLMs (Qwen2.5-VL 7B, MedGemma 4B, LLaVA-Med-v1.5-Mistral 7B); on the tuberculosis dataset MedGemma + FF reaches RadFact F1 0.3055 versus 0.2266 for MedGemma.
It designs a fully automated, LLM-bootstrapped data pipeline that builds a large-scale (image, multi-label) dataset from existing image-report pairs without manual annotation. Whereas prior label-guided methods such as TieNet assume a fixed vocabulary tightly coupled to a specific dataset, this pipeline automatically produces a unified taxonomy through batched extraction plus iterative hierarchical merging, with frequency-threshold filtering. Using GPT-5-mini with b=200, κ=200, θ=15, it yields 7 labels for tuberculosis and 42 for ophthalmology; on 50 randomly sampled ophthalmology reports, verified by a professional ophthalmologist, coverage and LLM-Match accuracy are both 100% (50/50), label redundancy is 7.5%, 80% of labels are valid and decidable, and none are invalid.
In the multi-label classification stage it adapts logit adjustment to multi-label binary classification to handle long-tail findings, where rare but critical findings appear in fewer than 1% of cases. Standard binary cross-entropy tends to overlook tail classes; this method shifts each raw logit by the log-odds of its empirical frequency (τ=1) before computing the loss, re-balancing the decision boundary. The guidance model reaches Macro-/Micro-F1 of 0.5227/0.7328 on tuberculosis and 0.7071/0.8426 on ophthalmology test sets, which the authors take as sufficient grounding.
Ablation analysis shows visual information and factual guidance are complementary, and identifies label quality as the key bottleneck. Five configurations on the tuberculosis dataset with Qwen2.5-VL show that image-only suffers mode collapse (precision 1.0000, recall 0.0145), adding predicted labels raises F1 to 0.2115, image plus predicted labels gives the best practical performance at F1 0.2831, while the oracle image + ground-truth labels reaches F1 0.4421 versus 0.3750 for labels-only ground truth. The conclusion comes from controlled configuration comparisons on the same dataset, covering predicted versus ground-truth label sources and image presence or absence.
Perspective
The framework targets clinical scenarios where reports revolve around targeted and enumerable finding categories, and the authors position it as plug-and-play and compatible with any MLLM architecture. It fits disease-focused datasets that already have image-report pairs but lack fine-grained finding labels, reducing manual annotation needs. For a reader, this means that in settings such as tuberculosis chest X-rays and multimodal ophthalmology (fundus, OCT, OCTA), one can first build a label taxonomy and multi-label classifier, then guide generation with label prompts.
No established clinical efficacy metric is currently available for the Chinese ophthalmology reports, so only NLG metrics are reported for that dataset and its clinical factual correctness awaits dedicated evaluation. The performance gap between predicted and ground-truth label conditions highlights label quality as the key bottleneck, and room for improving the taxonomy and classifier deserves attention. In addition, the tuberculosis dataset is small, and the weak zero-shot performance of closed-source models points to the necessity of domain-specific training; whether these results generalize to other diseases, other report languages, and larger data remains an open question.
