Public articles linked to the same research event.
arXiv The authors introduce a diagnostic benchmark with four task families covering structural grounding, synthesis-condition verification, evidence-sufficiency verification, and MLIP-based computational verification, evaluate T-MOF-1-3 under closed-book, retrieval-enabled, and oracle-evidence settings to localize failures in knowledge access, evidence acquisition, and reasoning, and build MOF-Verify, a failure-aware agentic harness that substantially improves MOF hypothesis-verification performance over direct inference and retrieval-based baselines across multiple backbone LLMs, with benchmark datasets released.
The authors introduce a diagnostic benchmark with four task families covering structural grounding, synthesis-condition verification, evidence-sufficiency verification, and MLIP-based computational verification, evaluate T-MOF-1-3 under closed-book, retrieval-enabled, and oracle-evidence settings to localize failures in knowledge access, evidence acquisition, and reasoning, and build MOF-Verify, a failure-aware agentic harness that substantially improves MOF hypothesis-verification performance over direct inference and retrieval-based baselines across multiple backbone LLMs, with benchmark datasets released.
The authors introduce a diagnostic benchmark with four task families covering structural grounding, synthesis-condition verification, evidence-sufficiency verification, and MLIP-based computational verification, evaluate T-MOF-1-3 under closed-book, retrieval-enabled, and oracle-evidence settings to localize failures in knowledge access, evidence acquisition, and reasoning, and build MOF-Verify, a failure-aware agentic harness that substantially improves MOF hypothesis-verification performance over direct inference and retrieval-based baselines across multiple backbone LLMs, with benchmark datasets released.
The authors introduce a diagnostic benchmark with four task families covering structural grounding, synthesis-condition verification, evidence-sufficiency verification, and MLIP-based computational verification, evaluate T-MOF-1-3 under closed-book, retrieval-enabled, and oracle-evidence settings to localize failures in knowledge access, evidence acquisition, and reasoning, and build MOF-Verify, a failure-aware agentic harness that substantially improves MOF hypothesis-verification performance over direct inference and retrieval-based baselines across multiple backbone LLMs, with benchmark datasets released.