Skip to main content
Back to timeline
arXivSource publication:

A hidden current date in system prompts shifts nine models' scores by up to 14% and reshuffles leaderboards

Synopsis

Holding all other settings fixed and varying only the hidden current date injected into the system prompt across every day of 2024, this study of 9 recent LLMs and 6 datasets finds accuracy swings of up to 6% on multiple-choice QA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, with model rankings reordering, an effect larger than batch size or numerical precision and one that chain-of-thought prompting amplifies rather than reduces.

AI-generated editorial illustration: Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

Interpretation

The study identifies a previously overlooked source of evaluation variance: API providers and some open-weight chat templates inject the current date into the system prompt at a layer users cannot see, and that information changes every day. Prior work attributed LLM evaluation instability to batch size, numerical precision, hardware, or prompt formatting; this work isolates time-varying hidden prompt metadata as a distinct source. The authors describe the injection as occurring at the template level or server-side, and cite Qwen3 reporting the current date when called through OpenRouter as evidence of server-side injection; they also note that not all 9 models inject the date in their default templates.

Changing only the date produces performance swings across tasks and model scales: up to 6% on multiple-choice QA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, with model rankings reordering. The swings appear on execution-graded code and on full-output translation metrics, showing the effect is not a multiple-choice artifact but a general property of evaluation under hidden time metadata. Coverage spans 9 models (Llama 3.1 8B/70B, Gemma 3 4B/27B, Qwen3 4B, Qwen3-Next 80B, Phi-4 14B, GPT-OSS 20B/120B) and 6 datasets (MMLU, GPQA, ARC-Challenge, GSM8K, HumanEval, WMT English to German, Finnish, and Czech), sweeping dates from January 1 to December 31, 2024; GSM8K averages 7.75% versus 2.52% on MCQA, HumanEval averages 4.81% with up to 7.32%, and MT averages up to 1.88 BLEU and 1.33 chrF.

There is no consistently good date to fix: the effect behaves as noise specific to each model and dataset, and it exceeds batch size, numerical precision, and GPU model while being comparable to option order and to rewording the system prompt. The authors place the date effect on the same scale as other known non-determinism sources using correlation and coefficient-of-variation measures, and test whether the effect is shared across models or datasets. Across 27 model-dataset combinations the Spearman correlation between date and accuracy has a median of 0.02; across-dataset and across-model correlations have medians of 0.03 and 0.01, and a date moves both accuracies in the same direction on 51% of dates, near chance; six system-prompt wordings yield an accuracy CV of 0.78%, the same as the date effect.

Chain-of-thought and few-shot prompting do not reduce date sensitivity, and chain-of-thought amplifies it; the authors recommend removing the date from the chat template when possible, or fixing it and reporting it alongside results. This contradicts the intuition that explicit reasoning or task examples ground the model and blunt the influence of superficial metadata, and it yields a concrete evaluation-protocol recommendation. Chain-of-thought on MMLU with Llama 3.1 (8B) across all 2024 dates shows the date affecting the selection of initial CoT tokens and cascading through autoregressive paths; with 5-shot prompting the gaps persist across all models, averaging a 2.27% accuracy delta versus 2.52% zero-shot.

Perspective

This work speaks to researchers and leaderboard maintainers who benchmark through default chat templates or API services, across four task types (multiple-choice QA, math reasoning, code generation, machine translation) and date values within 2024. Its next step is to keep the model's default template and special tokens but drop the current date where possible, or fix the date and report it with results when removal is not possible; it also suggests treating confidence and calibration as first-class evaluation criteria alongside accuracy to flag predictions prone to shifting under hidden factors.

The authors note in their limitations that experiments cover a selected set of models and datasets, and that the proprietary-model check is limited to GPT-5.1 over one week; the date is injected in a fixed format at a fixed position and varied only within 2024, so sensitivity to other formats, positions, or years remains to be studied, and open-ended tasks such as dialogue or summarization are not covered. A preliminary appendix experiment shows that varying the user location in the system prompt can also produce differences comparable in magnitude to the date effect, hinting that other hidden variable metadata may matter. In addition, many numeric cells in the result tables of the loaded text are empty, so per-model, per-dataset exact figures cannot be restated here; only the summary values stated in the prose are quoted.

Sources