Skip to main content
Back to timeline
arXivSource publication:

Across 52,500 judgments from seven open-weight models, repeated ratings proved highly correlated; a beta-binomial latent class model lifts 95% interval coverage to 0.90

Related research and updates

Synopsis

Using 52,500 judgments from seven open-weight models of 1B to 32B parameters, collected under three prompt conditions with ten samples per item on a public benchmark with human labels, the authors test the latent class assumption that repeated ratings are conditionally independent, find intraclass correlations of 0.50 to 0.93 on the balanced split so that ten ratings carry the information of 1.07 to 1.81 independent ratings, and propose a beta-binomial latent class model with a single correlation parameter, establishing its identifiability and that of extensions to covariates, treatment effects and partially labeled data, raising coverage of nominal 95% intervals for prevalence and the false acceptance rate to 0.90 on the balanced split.

AI-generated editorial illustration: Quantifying and Correcting Measurement Error in LLM-Generated Classifications: A Beta-Binomial Latent Class Model for Correlated Repeated Judgments

Interpretation

The paper directly tests the conditional independence assumption underlying latent class models for repeated LLM classification, reporting intraclass correlations between 0.50 and 0.93 on the balanced split of a public benchmark, so that ten ratings carry the information of only 1.07 to 1.81 independent ratings. Prior practice of aggregating repeated ratings into prevalence and error-rate estimates assumes ratings are conditionally independent given the true label; this work quantifies the departure from that assumption using 52,500 judgments from seven open-weight models of 1B to 32B parameters under three prompt conditions with ten samples per item. Evidence comes from 52,500 judgments with human labels on a public benchmark, spanning seven models and three prompt conditions, with interval estimates for intraclass correlation and equivalent number of independent ratings.

Under this correlation, a latent class model with a binomial likelihood is both miscalibrated and biased: its nominal 95% intervals for the false acceptance rate contained the value computed from human labels in none of 42 settings, and it underestimated the total error rate in all 42. This turns a violation of the conditional independence assumption from a theoretical concern into an observable inferential consequence, showing that both interval coverage and point estimation of conventional latent class modeling fail in this setting. The conclusion rests on direct comparison against values computed from human labels across 42 settings, using two checkable indicators: whether intervals contain the target value and whether the total error rate is underestimated.

The authors propose a beta-binomial latent class model with a single correlation parameter, establish its identifiability and that of extensions to covariates, treatment effects and partially labeled data, and raise coverage of nominal 95% intervals for prevalence and the false acceptance rate to 0.90 on the balanced split. Relative to the binomial-likelihood version, the new model absorbs correlation among repeated ratings through a single correlation parameter while providing identifiability proofs, giving explicit conditions for inference under correlated structure. Support comes from interval coverage results on the balanced split (0.90) and from simulation studies designed around the identification conditions.

Recovery of the latent state is governed by the gap between the model's acceptance rates on truly positive and truly negative items, which equals Youden's J and predicts recovery almost monotonically (Spearman correlation 0.96); when this gap is small, a labeled subset improves prevalence estimates but not the classification of individual items. This provides a single interpretable quantity for judging whether repeated ratings can recover latent labels, and distinguishes the different roles of labeled data at the population level versus the individual-item level. Evidence is the near-monotonic relationship with Spearman correlation 0.96, the observed difference in the effect of a labeled subset across the two tasks, and simulation studies designed around the identification conditions.

Perspective

The results are intended for research and engineering settings that estimate outcome prevalence and model error rates from repeated LLM ratings, applying to public-benchmark settings with human labels and multiple samples per item, and they extend to covariates, treatment effects and partially labeled data. For users, the work supplies a modeling framework for measurement-error correction under correlated structure and an operational basis for judging latent-state recovery via the gap between acceptance rates on truly positive and truly negative items (Youden's J); when that gap is small, a labeled subset can be used to improve prevalence estimates.

The intraclass correlation and coverage results are reported on the balanced split of a public benchmark, so behavior under other data distributions and tasks remains to be examined; identifiability conclusions depend on model specification and identification conditions, and whether these hold in a given application must be judged case by case; when the gap between acceptance rates on truly positive and truly negative items is small, there is limited room to improve individual-item classification, and the value of labeled data lies mainly in population estimates; the simulation studies are designed around the identification conditions, so their generalization to real annotation noise and more complex prompt conditions remains an open question.

Sources