Public articles linked to the same research event.
arXiv EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.
EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.
EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.
EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.