Skip to main content
Back to timeline
arXivSource publication:

Purdue team's covariate-dependent CR-MRF jointly models ordinal ratings and predicts pairwise preferences more accurately than Bradley-Terry and Plackett-Luce on MovieLens

Synopsis

The authors propose a covariate-dependent consecutive-ratio Markov random field (CR-MRF) that jointly models multivariate ordinal ratings, show that comparison models such as Bradley-Terry and Plackett-Luce arise as restrictions of it under independence and an exponential-race argument, and develop maximum likelihood inference with an importance-sampling estimator of the intractable normalizer; in simulations and on MovieLens, modeling ordinal scores directly yields more accurate and better-calibrated out-of-sample pairwise preference predictions than fitting comparison models to derived win-loss data.

Source-provided article image: Covariate-dependent Joint Modeling of Multivariate Ordinal Preferences and Its Connections with Comparison Models
Figure 1 ·

Figure 1: Observed and model-implied pairwise preference probabilities on the 64 held-out MovieLens reviewers, split by reviewer gender, with ordinal ratings ( M = 4 M=4 ).

arXiv

Interpretation

Introduces a covariate-dependent CR-MRF in which node and edge parameters are linear functions of covariates, enabling joint modeling of multiple ordinal attributes rather than separate per-attribute models. Prior CR-MRF work by Suggala et al. (2017) had no covariates and used pseudolikelihood; this work makes node and edge parameters functions of covariates and develops full-likelihood inference. Model definition and inference algorithm appear in Section 3, MLE existence conditions in Supplementary Section S.1, and convergence in Theorem 1.

Shows that the Bradley-Terry-Davidson model and the weak-ranking Plackett-Luce model can be recovered from the CR-MRF under independence restrictions via an exponential-race argument. The authors state this connection has not been noted in either the MRF or the comparison-model literature, clarifying that BT/PL amount to a coarsening of the ordinal observations. Derivations are given in Proposition 1, Theorem 2, and Corollary 1, with proofs in Supplementary Sections S.2.2-S.2.4.

Quantifies the Fisher information lost when ordinal ratings are converted to weak rankings: with at least three rating levels the difference is positive definite, so information is lost in every nonzero natural-parameter direction. Turns the informal claim that comparison models discard information into a provable Fisher-information inequality, holding even when all ties are retained. Proposition 2 states the result, with proof in Supplementary Section S.2.6 and explicit calculations for the exponential-race case in S.3.

In simulations and on MovieLens 1M, the CR-MRF that models ordinal scores directly outperforms covariate-dependent BT/PL in out-of-sample pairwise preference prediction. Prior work had not directly compared joint ordinal modeling against comparison models on out-of-sample preference prediction; this work reports RMSE and Brier score comparisons. Simulation experiments report RMSE and Brier scores (Table 1); the MovieLens application reports gender-stratified pairwise comparison probabilities (Figure 1), with supplementary results for random tie-breaking and binary recoding in S.6.2.

Perspective

The work targets settings where full ordinal feedback is available, such as MovieLens 1-5 star ratings or HelpSteer-style multi-attribute ordinal annotations, and assumes attributes are globally coupled in a way an undirected graph can capture. It is meant for readers who want to use covariates (e.g., gender, age) together with the joint structure of multiple attributes and who care about pairwise or list-wise preference prediction. The authors note that the simple node-conditional distributions allow missing data to be handled via conditional sampling, whereas comparison models require a strongly connected graph over items for maximum likelihood estimation. The conclusion proposes examining whether the CR-MRF loss could serve as an alternative to BT or PL loss in reward models for LLM alignment with human feedback.

The information-loss result in Proposition 2 is stated under finite natural parameters and at least three rating levels, so its scope should be read with those conditions. Supplementary Section S.3 notes that the Fisher orthogonality and efficiency ratio at equal rates are local statements, not uniform over all parameter values; away from equal rates the off-diagonal information generally does not vanish. In the MovieLens application, the authors report that CR-MRF improves more on ambiguous comparisons near 0.5, while PL-based estimates are better calibrated for movies with distinct preferences, so the advantage varies with comparison difficulty. In addition, the specific numeric values in the main-text tables are not fully rendered in the provided text; readers needing exact RMSE, Brier scores, and regression coefficients should consult the original tables and supplement.

Sources