Skip to main content
Back to timeline
arXivSource publication:

DUCX decomposes chest X-ray agent unfairness into tool exposure, tool transition, and reasoning, finding gaps up to about 50% beyond end-to-end metrics

Synopsis

Using MedRAX as the instantiated system, this work audits fairness in tool-using chest X-ray question-answering agents and proposes DUCX, a stage-wise decomposition that separates end-to-end bias into tool-exposure bias, tool-transition bias, and LLM reasoning bias; across five driver LLMs on CheXAgentBench and the curated MIMIC-FairnessVQA, demographic gaps persist end to end (equalized odds up to 20.79%, lowest fairness-utility tradeoff down to 28.65%), and subgroup disparities in tool usage, routing patterns, and reasoning traces are not predictable from end-to-end evaluation alone (for example, conditioned on segmentation-tool availability the subgroup utility gap reaches as high as 50%).

Source-provided article image: DUCX: Decomposing Unfairness in Tool-Using Chest X-Ray Agents
Fig. 1

Fig. 1. Overview of DUCX. (1) Dataset curation: we use CheXAgentBench and curate MIMIC-FairnessVQA into a standardized form. (2) MedRAX agent execution (ReAct): a driver LLM iteratively reasons, selects tools, and synthesizes a final answer, where we highlight multiple points where bias can be introduced (tool exposure, tool transi- tions, and final synthesis). (3) Fairness decomposition: We evaluate end-to-end fairness (ACC, ∆ACC, DP, EoD, FUT) and decompose it into tool exposure bias (gaps condi- tioned on tool presence), tool transition bias (gaps in tool routing), and LLM reasoning bias (gaps in synthesis behaviors).

· Page 3

Interpretation

It introduces DUCX, a stage-wise fairness decomposition that splits agent unfairness into tool-exposure bias (subgroup accuracy gaps conditioned on a tool being invoked), tool-transition bias (differences in subgroup tool-routing captured by Markov transition matrices), and LLM reasoning bias (subgroup gaps in JudgeGap, hedging, and demographic-term features of the final response text). Prior medical imaging fairness work largely treats a model as a single decision function and compares final predictions; this work extends auditing to multi-step agent pipelines so an observed disparity can be mapped to a specific stage. The framework is formally defined with equations and implemented on MedRAX's ReAct-style execution traces, with confidence intervals from non-parametric bootstrapping (1,000 resamples).

Across five driver LLMs (LLaMA3.1-8B, Ministral-3-8B, Qwen3VL-8B, Qwen3-8B, Gemini3-Flash) and two benchmarks, end-to-end fairness gaps persist: equalized odds up to 20.79% and the lowest fairness-utility tradeoff score down to 28.65%; Qwen3 reaches 69.10% and 49.89% accuracy on the two datasets while keeping relatively low ∆ACC and DP. The authors describe this as the first systematic demographic fairness evaluation of MedRAX-style chest X-ray agents under a unified setup, covering gender and age as sensitive attributes. Results are reported in a table with accuracy, ∆ACC, DP, EoD, and FUT plus confidence intervals, and the text notes that DP is closely tied to subgroup base rates and requires careful examination before adoption.

Intermediate behaviors show subgroup disparities that end-to-end evaluation does not predict: on CheXAgentBench segmentation has the largest and heaviest-tailed exposure-conditioned accuracy gaps with report generation a distant second, while classifier and grounding stay near zero; on MIMIC-FairnessVQA the largest disparities concentrate in the visualizer; conditioned on segmentation-tool availability the subgroup utility gap reaches as high as 50%. This indicates that the main fairness bottleneck shifts to different tools under different data distributions, and that overall ∆ACC is smaller than the worst exposure-conditioned gaps because it averages over heterogeneous trajectories. Presented as distributions of |∆ACC| across driver LLMs (violin plots), with each group's tool exposure rate also reported to avoid confounding with routing.

Tool routing and response synthesis also differ by subgroup: in both datasets female patients are more likely than male patients to proceed directly from the classifier or report generator, male patients in MIMIC-FairnessVQA often call the classifier again after the visualizer, and older individuals and male patients show more repeated grounding calls; reasoning bias is highly model-dependent, with Qwen3VL showing larger hedging gaps than all other backbones across both datasets and attributes, and low ∆Demo not implying low ∆Hedge or ∆JudgeGap. This shows that even with the same questions and tool outputs, driver LLMs can produce subgroup-dependent variation in synthesis behavior such as uncertainty expression, and that these measures need not move together with explicit demographic framing. Transition bias is shown as transition-matrix differences averaged across LLMs, and reasoning bias as subgroup gaps in three response-level features with confidence intervals.

Perspective

The work targets tool-using chest X-ray multiple-choice question-answering agents, instantiated as MedRAX's ReAct-style single-agent pipeline, with a tool pool of classifier, VQA, report generator, segmentator, visualizer, and phrase grounding, and with sensitive attributes limited to gender and age (threshold 60). Within this setting it provides a map from observed disparity to the responsible stage, which can be used to prioritize which tools and stages to audit or intervene on, and it releases MIMIC-FairnessVQA (400 studies, 2,000 instances with manually verified answers) plus open-source code so researchers can reuse the audit procedure on comparable agents. The authors state that future work will develop stage-targeted mitigation and extend audits to broader clinical tasks and agent designs.

Readers should still watch that the decomposition characterizes which stage a gap appears in rather than establishing causal chains between stages; that DP is closely tied to subgroup base rates, which the authors also flag as requiring careful examination; that reasoning bias is measured by an external LLM judge and pattern matching, so stability across judge models or matching rules remains to be seen; that MIMIC-FairnessVQA questions are LLM-generated from a MedRAX-style template with manual answer verification, leaving open how their difficulty distribution maps to real clinical questioning; and that although exposure analysis reports subgroup exposure rates to reduce confounding with routing, unobserved confounding between exposure and utility may remain.

Sources