Skip to main content
Back to timeline
arXivSource publication:

Dual-agent framework PsyCIDRA raises diagnostic agreement in free-form psychiatric interviewing: 60.5% vs 51.9% rank-1 on held-out simulated cases and 79.6% vs 65.4% agreement with psychologists on whether to propose a hypothesis in a 101-participant study

Synopsis

The authors present PsyCIDRA, a dual-agent framework in which an interviewer agent conducts free-form psychiatric interviews using working notes, expert-written skills, and ICD-11 retrieval, and a diagnostic reasoning agent then receives only the completed transcript and reports hypotheses with supporting, conflicting, and missing evidence, withholding a final hypothesis when evidence is insufficient; across a 53-case evaluation cohort and 81 held-out simulated cases it achieves higher diagnostic agreement than direct prompting, and in a blinded study of 101 human participants it agrees with psychologists on whether to propose a diagnostic hypothesis in 79.6% versus 65.4% of cases.

Source-provided article image: PsyCIDRA: A Dual-Agent Framework for Psychiatric Interviewing and Diagnostic Reasoning
Figure 1 ·

Figure 1: PsyCIDRA and its evaluation. Simulated patients use PsyCPG-generated profiles. PsyDARA analyzes only PsyCIRA’s visible transcript. Case diagnoses and expert judgments provide simulation and human references, respectively; neither agent sees these references.

arXiv

Interpretation

PsyCPG converts clinical cases into structured simulated-patient profiles with a typed fact ledger and disclosure-level controls. Relative to role-play directly from the clinical case, expert mean quality rose from 3.83 to 4.50 out of 5 and overall authenticity from 3.47 to 4.50, while first-diagnosis match to the intended ICD-11 code rose from 24/30 to 29/30. Two psychology experts each interviewed one condition on 30 paired cases, yielding 60 interviews under the same backbone and a 16-exchange limit, blinded to condition, source case, and reference diagnosis; the quality form is study-specific and has not undergone psychometric validation.

PsyCIDRA separates interviewing from diagnostic reasoning into two agents, with the interviewer using a private notepad, 60 expert-written skills, and ICD-11 retrieval tools, and the diagnoser receiving only the completed transcript and emitting a structured report. The report orders data sufficiency, problem representation, evidence states, alternative explanations, a differential with supporting and conflicting evidence, uncertainties, a separate safety assessment, and zero to three hypotheses with ICD-11 codes and qualitative confidence, leaving the final field empty when evidence is insufficient. Prompts, tools, skill catalogue, and report schema were jointly designed by NLP researchers and psychologists; tool-use counts are descriptive and cannot establish a relationship with diagnostic performance.

On simulated cases, complete PsyCIDRA achieves higher diagnostic agreement than direct prompting of the same backbone models. On the 53-case evaluation cohort it improves rank-1 accuracy by 7.5 to 15.1 percentage points and coverage by 5.7 to 18.9 points across four model settings; on the 81-case held-out cohort rank-1 is 60.5% versus 51.9% and coverage 70.4% versus 60.5%. The evaluation cohort spans four backbones and four interviewer/diagnoser combinations, and the held-out cohort is disjoint from evaluation; both cohorts fall below evaluation-cohort levels, and 44/81 held-out cases fall outside the anxiety/fear, OC-related, and mood groups.

In a blinded feasibility study with 101 human participants, PsyCIDRA agrees more often with psychologists on whether to propose a diagnostic hypothesis. Agreement is 79.6% versus 65.4%; among transcripts with an expert-proposed diagnosis, reference-code coverage is 90.0% versus 82.6%, with the larger separation coming from withholding when experts also withhold. Two psychologists independently reviewed all 101 transcripts and reconciled disagreements, blinded to interview mode and model-generated reports; complete-system comparisons involve different participants, limiting attribution of differences to the systems.

Perspective

The work addresses researchers and clinical-informatics teams building or evaluating psychiatric interview agents, in settings where the system produces an interview transcript and a structured report for qualified experts to inspect rather than delivering diagnoses to participants. PsyCPG's inputs separate the source case and diagnostic reference from cohort-level personality settings, an expert-authored sociocultural catalogue, source-grounded safety constraints, and scheduled interaction challenges, which supports reuse with other case collections after adapting those inputs and validating the resulting patients. The held-out test and human study used Terra configurations frozen after evaluation, and the 16-exchange interview cap was chosen to limit participant burden, with a 24-exchange cap giving no consistent rank-1 benefit.

Several open questions remain for a careful reader: both experts and diagnostic agents relied on information elicited by the interviewer, and the experts did not conduct independent clinical interviews, obtain collateral histories, or assess participants longitudinally, so agreement measures consistency with expert interpretation of the available transcript and shared omissions may produce agreement without establishing diagnostic correctness. Text-only interaction excludes direct observation of appearance, facial expression, psychomotor behavior, and vocal characteristics. The human cohort consisted of self-selected adult volunteers in Iran, mostly aged 18 to 34 and holding a bachelor's or master's degree, and cannot represent the general population or a clinically stratified sample. Simulation provides controlled comparisons but does not replace evaluation with real patients. Structured reports do not establish faithful explanations or clinical benefit, and the study did not evaluate improvements in clinicians' decisions, review time, or patient outcomes. On safety, AISC flags occurred in 2/52 Direct and 3/49 PsyCIRA transcripts, including two PsyCIRA crisis-handling concerns, and one analyzed session in each arm activated the live safety stop; these counts describe the analyzed cohort only. PsyDARA's confidence labels are ordinal rather than calibrated probabilities, its high-confidence category contains only five reports, and usefulness for expert review remains to be tested.

Sources