Skip to main content
Back to timeline
arXivSource publication:

Asked the same prompt twice, eight models returned identical structure node sets in only 28% of cells, and the authors' self-audit withdrew six of seven claims

Synopsis

The authors ran an end-to-end self-audit of an evaluation of LLM-inferred prompt structure: across eight open model variants spanning five families and 8B to 675B parameters, with caching disabled and 293 raw intermediate representations persisted, repeated identical calls yielded mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells were never node-set-perfect; a joint cluster bootstrap showed only the bottom of the ranking is firm (99% and 86% rank retention), the middle four at 27% to 48% and the top two at 68% each; two equally defensible merge rules changed four of eight rows and moved the study-wide headline by 7 percentage points; reproducibility could not be read as accuracy; and four of eight endpoints were withdrawn within ten weeks, so the study as specified can

AI-generated editorial illustration: How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure

Interpretation

On prompt-structure inference, identical calls do not reliably return identical structure: mean node-set Jaccard runs from 0.39 to 0.96 and ID-stability from 0.22 to 0.96, only 35 of 127 prompt-model cells (28%) are node-set-perfect, and 64% to 100% of nodes in the synthesised corpus are explicit annotations the model was told to copy. Earlier work on non-determinism under nominally deterministic settings was largely general or aimed at structured extraction; this work measures it specifically for prompt-structure inference across model families and parameter scales and releases the raw outputs. Eight model variants across five families from 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted; means carry 95% percentile bootstrap intervals over 10,000 prompt resamples, for instance gpt-oss:120b at 0.39 with a wide interval and minimax-m2.7 at 0.70 with an interval that is also not narrow.

Auditing the ranking itself with a joint cluster bootstrap leaves only the bottom firm: gpt-oss:120b stays last in 99% of replicates and qwen3-coder-next seventh in 86%, the middle four range from 27% to 48%, and the top two hold position in 68% each, so the table reliably identifies the worst model but not the best. The authors turn the uncertainty analysis normally applied to models back onto their own leaderboard, and state explicitly that they introduce no new rank-uncertainty estimator; the contribution is applying uncertainty analysis as one stage of an end-to-end audit. Of 10,000 replicates, 9,583 were usable; one prompt multiset is drawn per replicate and every model is scored on that same draw, preserving the pairing; the authors note retention frequency is a bootstrap diagnostic, not a posterior probability.

Two equally defensible rules for merging repeated campaigns change four of eight rows and move the study-wide headline by 7.1 percentage points; separately, the stabilisation criterion is satisfied by construction at its endpoint (48 of 98 cells) and the tolerance was documented as relative but implemented as an absolute 0.05 nodes, giving 80 versus 50 of 98 cells stabilising under the two readings. These failures were found without new data, by executing sensitivity comparisons rather than asserting them, which grounds the recommendation to publish per-cell provenance, execute sensitivity comparisons, and state the merge rule and every tolerance's units. The merge-rule sensitivity is reported as a full table (node-set-perfect cells moving from 35/127 to 26/127); the stabilisation section reports full counts under both tolerance readings and states that nothing in the data distinguishes them.

Reproducibility cannot be read as accuracy: re-reading the 293 persisted intermediate representations offline gives mean recall of annotated node ids of 0.86 and mean precision restricted to nodes the model marked explicit of 0.82, so roughly one in five nodes a model asserted it had copied corresponds to no annotation; stability and correctness also order the models differently, and none of the four correlations survives removal of the single low outlier. This separates run-to-run agreement from agreement with truth and gives a concrete consequence: a practitioner choosing on the table alone would take mistral-large-3:675b over ministral-3:8b and get the less accurate of the two. Correctness comes from offline re-reading of the annotated synthesised prompts with a deterministic parser, requiring no new inference; correlations are computed over the seven models measured both ways and reported with the drop-one result.

Perspective

This work addresses evaluators and tool builders who use hosted endpoints and build leaderboards from tens of prompts and a handful of models, in the setting of prompt-structure inference, where a program-style toolchain must first recover a prompt's structure. It offers a set of procedural safeguards: treat raw per-run outputs as the primary artifact, date every measurement, present the ranking as a dated document, report rank stability beside any ranking, publish per-cell provenance, and execute sensitivity comparisons rather than asserting them. The authors also note that the drift, stabilisation and correctness analyses regenerate offline only because the 293 raw intermediate representations were persisted; had only summary CSVs been kept, much of the paper would be unfalsifiable.

The main table is an incomplete, unbalanced panel: campaigns covered different prompt subsets and completion failures further changed each row's support, so raw cross-row differences are descriptive rather than controlled pairwise effects; two rows rest on 3 and 6 prompts, no multiplicity correction is made for the 28 implied comparisons, and the authors suggest reading a sorted table as a partition into clearly reproducible, clearly not, and undetermined. Between the recovery campaign and the annotated campaign, annotation status, corpus, prompt content, model panel and completion pattern all change together, and only four models appear in both panels, so the cross-regime comparison is a level difference rather than an isolated effect of recovery. The metadata bucket in the drift decomposition measures per-node metadata disagreement only, because the intermediate representation as emitted is a flat node list. The accessibility finding has been withdrawn by the authors: the timeout is the leading candidate but the cause is unresolved and the historical failure was not reproduced. In addition, the loaded text is the full paper, but some interval values appear in the prose as ranges rather than as enumerated endpoints, so exact figures still require the original tables.

Sources