According to research published on arxiv.org, researchers have developed a model-agnostic framework that allows frozen large language models (LLMs) and vision-language models (VLMs) to learn from deployment experience without requiring fine-tuning or access to model weights.
The framework addresses a significant limitation in medical AI: models typically remain frozen after deployment and “do not learn from the cases they solve,” according to the paper. This is particularly concerning in medicine, where “new clinical evidence, updated guidelines, and new therapies can change established practice.”
The system incorporates three forms of external expertise: a Skill component that guides reasoning and tool use, a Knowledge Memory storing reliable facts from earlier cases or trusted evidence, and a Multimodal Knowledge Base preserving visual examples. According to the researchers, a validation strategy “keeps an update only if it helps on new cases without degrading performance on earlier ones.”
Tested across six benchmarks covering clinical diagnosis, clinical workflows, medical reasoning, and visual reasoning tasks, the framework improved performance during online deployment by up to 34.2% over base models on medical tasks, according to arxiv.org. The researchers report the framework “generalizes to unseen cases, transfers to other models without further optimization, and works in non-medical domains.”
The testing included four open-weight and closed-source base models, demonstrating the approach’s model-agnostic capabilities.