Skip to main content
Back to timeline
arXivSource publication:

Frozen models accumulate external expertise during deployment: up to 34.2% online gain on medical tasks across six benchmarks, with zero-step transfer to other models

Synopsis

The work presents a model-agnostic framework that lets frozen LLMs and VLMs keep learning during deployment through three forms of external expertise: a Skill guiding reasoning and tool use, a Knowledge Memory storing reliable facts, and a Multimodal Knowledge Base that keeps visual cases and guides the model to compare each retrieved case with the current image; an update is kept only if it helps on new cases without degrading earlier ones. Across six benchmarks and four base models, the framework improves online medical performance by up to 34.2% over the base model, generalizes to unseen cases, and transfers to other models without further optimization.

AI-generated editorial illustration: Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI

Interpretation

The framework lets frozen models improve continually during deployment without access to model weights, so it applies to both open-weight and closed-source models. Unlike fine-tuning or reinforcement learning that update parameters, it turns experience into reusable expertise outside the model; unlike parameter-free methods such as EvoSkill and SkillOpt, it does not select updates on a fixed validation set. Across six benchmarks and four base models (Qwen3.8-27B, Nemotron-3-Nano-Omni-30B, GPT-5.6-Sol, GPT-4o-mini), it scores above Base in all 24 benchmark-model pairs and achieves the best score in 22, with an average gain of 17.6% versus 5.6% for ACE.

The validation strategy is necessary for the gains: a candidate update is accepted only if it improves on new cases without lowering the average on past cases. This replaces repeated selection on a fixed validation set, and periodically compares the current version with a retained earlier version, restoring the earlier Skill and Memory when the current version performs worse. In the ablation, accepting every update without validation lowers all six benchmarks by 21.8% to 37.8% relative to Full and leaves every score below Base; under five random case orders, average gains on AgentClinic and QCalEval are 5.7% and 16.5%, each at least five times the largest standard deviation.

MMKB preserves images, source answers, and available explanations, and guides the model to analyze how each retrieved case relates to the current image rather than only supplying retrieved content as context. Compared with text-only experience storage and with multimodal retrieval systems that mainly provide retrieved content, MMKB explicitly asks the model to compare observations and conditions and to consider differences, missing evidence, and contradictions. With identical retrieved references, standard multimodal retrieval changes the score negatively relative to Base, whereas MMKB improves all four image-based tasks, by up to 23.9% on MedChain Task 3; removing MMKB drops QCalEval by 14.0% and VisuLogic by 8.2%, returning both to Base levels.

Learned expertise generalizes to unseen cases after being frozen and transfers to other models with zero further learning. Because expertise is kept outside the model, a new or upgraded base model can use it immediately instead of learning it again. In frozen evaluation, Ours scores above Base in 21 of 24 pairs and best in 19, with an average gain of 20.7% (SkillOpt 8.3%, ACE 1.9%); cross-model transfer improves 33 of 36 results with an average gain of 20.0%, including expertise learned with open-weight Qwen3.8-27B raising closed-source GPT-5.6-Sol by 11.8% on AgentClinic and 23.2% on QCalEval.

Perspective

The framework targets streaming settings where source answers or reference reports become available during deployment: cases are answered first, and their source answers and scores are released only after the batch, so feedback can affect only later batches. It suits teams that want to improve closed-source or open-weight models without weight access, and clinical or multimodal tasks where better diagnostic procedures, reliable medical facts, and useful visual patterns can be reused. Expertise lives outside the model, so it can be carried over when the base model changes; gains are also observed on non-medical visual reasoning tasks.

Several open questions remain for a careful reader: online gains depend on case order and batch composition, and although five random orders show stable gains, matching seeds does not guarantee deterministic model responses; the MedChain implementation strictly follows the sequential workflow in which each task receives only the model's outputs from preceding tasks, so its scores are not directly comparable to those in the original paper's table; QCalEval's initial references are not uniformly expert-authored question-answer pairs, and 12 source-data references plus one figure reference come from the same study; in addition, this material is a full-text parse in which figures and some table contents are not itemized, so specific numeric details should be checked against the original figures and tables.

Sources