Skip to main content
Back to timeline
arXivSource publication:

How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards

Synopsis

The work introduces ECtHR-NPD, a benchmark of 14,575 European Court of Human Rights judgments with case-level non-pecuniary damage (NPD) awards in nominal euros, chronological splits and ID/OOD/Challenging diagnostic views, and evaluates six method families (constant predictors, gradient-boosted trees, retrieval methods, fine-tuned encoder LMs, prompted decoder LMs and knowledge-augmented agents), finding that more sophisticated LM and agentic approaches do not consistently outperform the strongest feature-based baseline and that all families struggle to identify zero awards and to calibrate high-award predictions.

Source-provided article image: How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards
Figure 1 ·

Figure 1: The ECtHR-NPD task. Models receive case metadata, violated Convention articles, case facts, and external macroeconomic covariates. Award-related material is excluded from model input. The model outputs the case-level non-pecuniary damage award.

arXiv

Interpretation

It formulates continuous non-pecuniary damage award prediction as a regression task, extending legal NLP evaluation to discretionary monetary remedies that lack a statutory formula. Prior legal benchmarks use largely categorical, retrieval-based, textual or rule-bounded targets such as sentencing terms or tax liabilities, whereas this work targets ECtHR Article 41 NPD amounts as a continuous outcome. The task is precisely defined (inputs are case metadata, the Facts section, violated articles and macroeconomic covariates; output is a non-negative euro amount), with a legal-expert-informed target construction and filtering procedure.

It releases ECtHR-NPD: 14,575 validated case-level targets, chronological splits, three diagnostic test views and structured annotations. Data come from public HUDOC judgments and pass eight validation checks (head separation, per-applicant sum consistency, currency normalisation, recoverability from operative provisions, among others), with target-construction material explicitly separated from model input. An exclusion cascade from 18,367 English judgments yields 14,575 cases; train 10,217, validation 1,461, test 2,897; zero-award rationales are quantified (finding_sufficient 58.2% of zeros, no_claim 36.1%).

It compares six method families with diagnostic evaluation, showing that stronger LMs and agentic configurations do not reliably beat feature-based baselines and that all families fail at zero-award recognition and high-award calibration. It provides error decompositions by award range, test view, respondent state and violated article, exposing failure modes that classification-oriented legal benchmarks do not measure. CatBoost attains the lowest aggregate MAE (EUR 9,881), 10.2% below the training-set median; paired bootstrap tests show CatBoost significantly better than the train median and the strongest prompted LM but not significantly better than BGE-M3 dense or ModernBERT late fusion; Challenging-view MAE is roughly twice full-test MAE.

Perspective

The benchmark targets case-level Article 41 non-pecuniary damage amounts in English-language ECtHR judgments, suited to studying continuous monetary remedy modelling and calibration evaluation in legal NLP; targets are nominal euro amounts, so it measures fit to the Court's realised practice rather than normatively just outcomes. Data and code are released so researchers can reproduce document retrieval and compare heterogeneous systems under one protocol.

A careful reader would still watch: residual label noise may remain in older judgments and multi-applicant bundled-award cases; Article 41 awards are equitable, so several amounts may be legally defensible on the same facts while point-error metrics treat any deviation from the realised award as wrong; prompted-LM and ReAct agent results are not averaged over repeated runs and may vary run to run; and the post-2021 test window contains a disproportionate share of Russia- and Ukraine-related cases, a distribution shift. This parse loaded the full text, so figure and appendix table details should be checked against the original.

Sources