Skip to main content
Back to timeline
arXivSource publication:

Porcedda's Sys1Cal-v1 shows Jev's Choice probabilities are systematically distorted, and restoring an uncertainty component lifts median soft accuracy from 0.771 to 0.978

Synopsis

Porcedda built Sys1Cal-v1, a synthetic dataset of 365 True/False items derived from 92 probability problems whose proposition probabilities are known by construction, and used total variation distance and distributional overlap (soft accuracy) to evaluate Jev's Noul, Choice, and Score primitives plus the open-source baseline SemIf, finding that Jev's Choice probabilities are systematically distorted and that inverting a latent third truth value of uncertainty estimated from Score expectations raises median Choice soft accuracy from 0.771 to 0.978.

AI-generated editorial illustration: Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration

Interpretation

The paper introduces Sys1Cal-v1: each item defines a proposition whose true probability is randomly generated, then renders it across six state types (explicit probabilities, frequencies from counts, compound probability, conditional probability, Bayes' rule, sequential Bayesian updates) and multiple equivalent representations (direct statements, counts, ratios, tables, prose, nested state, distractor-augmented state), totaling 365 rendered examples from 92 problems. Available public benchmarks evaluate confidence calibration, that is, whether predictions are overconfident or underconfident, not whether every returned option probability has the right numerical meaning; because Sys1Cal-v1 knows the true probability by construction, it compares the returned distribution itself. Dataset size and construction are stated in the text (365 rendered examples, 92 problems, six state types, multiple equivalent renderings) with a simplified item example; evaluation uses total variation distance and the distributional overlap it defines (soft accuracy).

On the same items, Jev's three primitives calibrate differently: Noul averages OVL 0.918, Score 0.886, and Choice 0.764, while the open-source Choice-style baseline SemIf reaches 0.629; Jev-Choice tends to return very high probabilities when the true probability is low. No public test had examined Jev's probability calibration, nor compared whether multiple probability interfaces of the same model share compatible semantics on the same items; this result makes the phenomenon of differing probability semantics across primitives of one model explicit. Each item is input 10 times to estimate returned probabilities, and a table reports mean TV and mean OVL per primitive; Score expectation aligns closely with Noul, indicating that the Equation 1 projection recovers Noul's probabilities.

The paper finds Score expectations are linearly related to Noul probabilities and close to ground truth, but non-linearly related to Choice probabilities, and therefore proposes a latent three-component distribution (True, Uncertain, False), estimates the uncertainty component from Score expectations, and fits a Score-to-Choice distortion map; introducing this component reduces the mean absolute error in predicting binary Choice probabilities, with a 95% confidence interval excluding zero. This reframes the Choice deviation from mere miscalibration into a collapse of a richer state into a forced binary output, and provides an estimable latent structure rather than only reporting error magnitude. Global and problem-level scale parameters are estimated across the 92 latent problems (global 0.978, 95% CI [0.789, 1.000]; latent 0.821, 95% CI [0.760, 0.882]), with goodness of fit and a confidence interval for the error reduction; the authors explicitly state this is an observational account, not a causal claim about Jev's internal implementation.

The uncertainty model has two operational uses: inverting the map yields corrected binary Choice probabilities, raising mean OVL from 0.764 to 0.880 and median OVL from 0.771 to 0.903; keeping the uncertainty mass yields a T/U/F interval representation with median interval OVL 0.978 and mean uncertainty width 0.116, and in a selective decision example the preferred action shifts from act True to act False to defer. The correction layer improves probability recovery while preserving the original Choice interface, while the T/U/F representation additionally exposes unresolved probability mass, connecting calibration to selective prediction and abstention. Tables report mean and median OVL for raw Choice, calibrated Choice, and T/U/F, plus a decision example with concrete loss values; the authors caution that interval OVL is not directly interchangeable with binary OVL and must be reported with interval width.

Perspective

Sys1Cal-v1 is designed to isolate probability semantics, not to measure broad natural-language competence; its synthetic construction gives each item an exact pointwise probability, which is what makes distributional overlap meaningful in this setting. The benchmark targets structured decision models that treat probabilities as first-class outputs, and applies to cost-sensitive classification, autonomous decision systems, and settings where a model may abstain or defer. The two uses map to different deployment patterns: when an application needs a single calibrated binary probability for thresholding, ranking, expected-utility decisions, or risk scoring, inverse uncertainty calibration applies; when the system can expose richer state, the T/U/F representation supports acting automatically when uncertainty mass is small and requesting more evidence or routing to a slower model or human reviewer when the interval crosses a decision threshold.

The authors note the current release uses controlled probability families and templated renderings, and that future versions should include richer linguistic variation, adversarial paraphrases, domain-specific decision problems, and independently validated natural-language formulations. The Score analysis relies on a one-dimensional projection, the expectation of a 10-level ordered distribution; other summaries, or direct evaluation of the full ordered distribution, may expose additional structure. The uncertainty model is explicitly framed as an observational account of the relation between Score and Choice, not a causal claim about Jev's internal implementation: the results show Choice behaves as if a richer state were collapsed into a forced binary output and that the inferred unresolved mass is useful for calibration and decision making, but they indicate rather than prove that Jev explicitly represents a hidden third truth value internally. In addition, representation sensitivity can hide behind aggregate calibration, since a model may achieve a reasonable average error while still changing its probabilities across equivalent formulations, so representation sensitivity and interval width should be read alongside the headline metrics.

Sources