Skip to main content
Back to timeline
Journal of Chemical Information and ModelingSource publication:

In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts

Synopsis

The study evaluates nine LLM variants across three families (GPT-4.1, GPT-5, Gemini 2.5) on three MoleculeNet datasets (Delaney solubility, Lipophilicity, QM7 atomization energy) using a systematic blinding procedure that iteratively reduces available information, complemented by 0-, 60-, and 1000-shot in-context sample sizes, in order to determine whether molecular property prediction reflects genuine in-context regression or verbatim retrieval of memorized target values, and it adds positive and negative controls for the memorization experiments, structural reference baselines for the multi-shot experiments, and bootstrap confidence intervals for all results; it finds no evidence of verbatim retrieval on these legacy benchmarks and shows that blinding exposes conflicts between pre-traine

AI-generated editorial illustration: In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts

Interpretation

The paper proposes and implements a systematic blinding evaluation framework that iteratively reduces available information to test whether LLM molecular property prediction relies on memorization. Prior evaluations of molecular property prediction typically report predictive accuracy alone, whereas this work treats information availability itself as a controlled variable, using progressive blinding to separate genuine in-context regression from memorized retrieval. The abstract states that the framework spans three MoleculeNet datasets and nine LLM variants, with positive and negative controls and bootstrap confidence intervals, making it a systematic methodological design.

On the three legacy benchmarks Delaney solubility, Lipophilicity, and QM7 atomization energy, no evidence of verbatim retrieval of memorized target values was found. This directly addresses concerns that widely used benchmarks may suffer from training data contamination, turning the question of whether contamination is actually exploited from speculation into a testable one. The conclusion rests on a positive and a negative control for the memorization experiments and on bootstrap confidence intervals for all results, making it a controlled empirical test.

The blinding experiments reveal conflicts between pre-trained knowledge and in-context information. This moves the focus from whether a model can predict to what happens when internal model knowledge disagrees with information supplied in the prompt, offering a new angle on the interplay between in-context learning and pre-trained knowledge. The evidence comes from a blinding sequence that progressively reduces information, with 0-, 60-, and 1000-shot in-context sample sizes serving as an additional control for information access.

The paper provides a principled framework for evaluating molecular property prediction under controlled information access. The framework integrates memorization testing, blinding design, and multi-shot controls, and can be reused by later work on other datasets or model families. The abstract explicitly states that the framework includes positive and negative controls for the memorization experiments, structural reference baselines for the multi-shot experiments, and bootstrap confidence intervals, forming a reproducible evaluation protocol.

Perspective

The framework targets evaluation of LLM molecular property prediction under controlled information access, and is suited to researchers who want to distinguish memorized retrieval from genuine in-context regression and to examine conflicts between pre-trained knowledge and in-context information; its conclusions rest on the three MoleculeNet datasets Delaney solubility, Lipophilicity, and QM7 atomization energy, and on nine variants across the GPT-4.1, GPT-5, and Gemini 2.5 families.

The text currently available is the abstract and metadata of the paper, without specific numerical results, figures, or experimental details, so the error levels on each dataset, the behavior at each blinding stage, and the concrete form of the knowledge conflicts remain open questions; whether the conclusions extend to other datasets, other model families, or other molecular properties also awaits follow-up work.

Sources