Skip to main content
Back to timeline
arXivSource publication:

Adapting LUH uncertainty heads to Persian medical models yields single-pass claim-level hallucination detection with PR-AUCs of 0.4820 and 0.4652

Related research and updates

Synopsis

The study adapts the LLM Uncertainty Head framework to two Persian medical models built on Aya-Expanse-8B, Gaokerena-V and Gaokerena-R; it first observes on a 168-question Iranian medical entrance examination that Gaokerena-V has substantially lower five-run consistency than Aya-Expanse-8B while Gaokerena-R is comparable, then builds Persian claim-level hallucination datasets with 1,600 responses per backbone and trains lightweight claim-level heads on frozen backbone attention maps and token probabilities, obtaining held-out PR-AUCs of 0.4820 and 0.4652 (2.30 and 2.66 times their random baselines) and ROC-AUCs of 0.7852 and 0.7810, with no retrieval or repeated sampling at inference.

Source-provided article image: Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models
arXiv · Page 4

Interpretation

The work transfers the LLM Uncertainty Head (LUH) framework to Persian medical models based on Aya-Expanse-8B, using Gaokerena-V and Gaokerena-R as two previously developed backbones. The text notes that existing uncertainty-head resources do not directly transfer to a new backbone and language and that repeated-sampling approaches are expensive; this work rebuilds that adaptation path for the Persian medical setting. Evidence comes from the adaptation procedure and the two-backbone experimental setup described in the abstract; it is an initial study with small test splits and automatically generated labels.

On a 168-question Iranian medical entrance examination, Gaokerena-V shows substantially lower five-run consistency than Aya-Expanse-8B, whereas Gaokerena-R is comparable to Aya-Expanse-8B. This observation treats response variability as a backbone-level prior signal, motivating and providing a contrast for the subsequent claim-level uncertainty modeling. Based on a five-run response-consistency comparison over 168 questions, reported directly in the abstract.

The study constructs two paired claim-level hallucination datasets directly in Persian, containing 1,600 responses for each backbone, and trains lightweight claim-level heads on frozen backbone attention maps and token probabilities. Unlike approaches relying on retrieval or repeated sampling, this design compresses uncertainty estimation onto internal signals available in a single forward pass. Dataset size and training-signal sources are stated in the abstract; labels are automatically generated.

On held-out test splits, the two heads obtain PR-AUCs of 0.4820 and 0.4652, corresponding to 2.30 and 2.66 times their respective random baselines, and ROC-AUCs of 0.7852 and 0.7810, while requiring neither retrieval nor repeated sampling at inference time. The result provides initial quantified performance for single-pass claim-level uncertainty estimation on Persian medical models and shows the improvement over random baselines. Metrics come from held-out test splits; the text also states that the test splits are small and the labels automatically generated, so this is initial evidence.

Perspective

The work targets Persian medical question answering, applies to models built on Aya-Expanse-8B with Gaokerena-V and Gaokerena-R as backbones, and produces claim-level uncertainty in a single forward pass at inference without retrieval or repeated sampling. For developers and researchers who want claim-level hallucination signals at low inference cost, this setting offers a directly comparable starting point: the five-run consistency observation on the 168-question Iranian medical entrance examination can be used to gauge backbone response stability, while the Persian claim-level datasets of 1,600 responses per backbone and the lightweight heads trained on frozen backbone attention maps and token probabilities form a reusable adaptation pattern. The text positions itself as an initial study of single-pass claim-level uncertainty estimation for Persian medical language models, so its conclusions apply to this language, domain, and backbone combination under the reported experimental setting.

The text states that the test splits are small and the labels automatically generated, so the stability of PR-AUCs 0.4820 and 0.4652 and ROC-AUCs 0.7852 and 0.7810 and the influence of label noise remain open questions. The abstract does not give the head architecture, training details, dataset construction procedure, or automatic labeling method, nor does it report confidence intervals or statistical comparisons across backbones; these require the full text. In addition, the relationship between the 168-question consistency observation and the claim-level datasets, and how the adaptation performs beyond Persian medical language and these backbones, remain to be examined.

Sources