Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

ClinCoT pushes preference optimization from answer-level correction down to lesion-region reasoning: it beats MMedPO and other baselines on most metrics across SLAKE, VQA-RAD and IU-Xray, with more consistent gains after SFT initialization

The work proposes ClinCoT, a clinical-aware visual chain-of-thought framework: it uses hypotheses-driven region proposals (disease-conditioned activation maps from a clinical-aware tool such as MedKLIP, thresholded into candidate regions), has the target Med-VLM generate multiple region-conditioned reasoning chains by jointly processing the original image and each candidate region, scores them with two Med-LLM evaluators that combine current-step score with next-step influence and penalize evaluator disagreement, and trains with a score-margin-aware DPO plus iterative regeneration of preference data; across SLAKE, VQA-RAD and IU-Xray it outperforms DPO, Self-Rewarding, STLLaVA-Med, POVID, SIMA, FiSAO and MMedPO on most metrics, is strongest on report generation, falls below MMedPO on VQA-R