Skip to main content
Back to timeline
bioRxivSource publication:

BoneGraph: A Domain-Specialised, Self-Correcting Reasoning System for Bone Science

Synopsis

This work builds BoneGraph, a bone-science-specific system delivered as a five-tab web application over a shared substrate of 7,449 documents, 248,629 SPECTER2 passage vectors, and a bone knowledge graph of 1,597 concepts and 1,699 causal relations, reporting MRR 0.928 on a 30-question seven-domain retrieval benchmark, 92.6% accuracy for the bone-region classifier on held-out MURA, and an increase in answer accuracy from 42% to 78% when the correct passage is supplied, with all inference performed locally on a single NVIDIA Jetson AGX Orin and no third-party API calls.

AI-generated editorial illustration: BoneGraph: A Domain-Specialised, Self-Correcting Reasoning System for Bone Science Retrieval, Grounded Inference, and Image Mechanics

Interpretation

It constructs a bone-science-specific retrieval substrate: 54,634 paper metadata records collected through OpenAlex using a 133-query, 17-group keyword taxonomy, a three-tier PDF resolution pipeline yielding 7,674 validated full-text PDFs plus 16 open-access textbooks, and a processing chain of PyMuPDF extraction, two-pass body-text language filtering, and sentence-aware chunking that produces 248,629 passages embedded with SPECTER2 (proximity adapter for documents, ad-hoc query adapter for queries) into 768-dimensional vectors. Existing open scholarly indexes (OpenAlex, PubMed) and citation-informed scientific embeddings (SPECTER, SPECTER2) are not curated for bone science; this work combines them and adds a bone-specific keyword taxonomy and a PDF-resolution stage to produce domain-focused retrieval. The corpus and pipeline are reported step by step in the methods (54,634 metadata records, 7,674 PDFs, 7,433 English full texts, 248,629 chunks), and retrieval is supported by a manually constructed 30-question, seven-domain benchmark with MRR 0.928, lowest in simulation (0.833) and biomaterials (0.867).

It implements a self-correcting reasoning loop: a reasoning agent answers only from the question, conversation history, and user rules without access to retrieved evidence; a deterministic check of eight built-in physical-grounding rules written in pure Python (cortical modulus, trabecular modulus, cortical density, trabecular BV/TV, WHO T-score, Wolff's-law direction, density-strength power-law scaling, lytic lesion effect) then screens the draft; and a critic holding retrieved literature and one-hop knowledge-graph edges returns accept, dispute, or conflicting_evidence, with the loop hard capped at two critic iterations. Where self-refinement lets a model judge its own output, this work substitutes a deterministic physical check and an evidence-anchored critic; where self-consistency relies on majority voting, the authors note that this presupposes a single extractable answer and does not fit the open-ended prose of bone science. The rule kinds (range, forbid pattern, comparative) and the specific thresholds of all eight rules are listed individually in Table 10; the constraint that the critic receives only raw one-hop edges and never pre-composed multi-hop chains is motivated by an audit on five fracture-domain probes, in which the literature channel returned on-topic passages at cosine similarity 0.78 to 0.84.

User feedback is compiled into durable, inspectable, reversible rules: after a user downvotes an answer and supplies free-text correction, llama3.2:3b compiles it into a structured editable rule that the user confirms into a personal registry, after which every future request both primes the agent with the rule and checks it deterministically in grounding; a worked walkthrough shows an erroneous premise about the failure order of cortical versus trabecular bone in osteoporosis being corrected once the rule is in effect, with no change to model weights. General-purpose LLMs offer no durable mechanism to accept an expert correction, so the same mistake may recur on the next similar question; this work places learning in retrieval and rule application that can be listed, edited, toggled, and deleted rather than in opaque weight updates. The walkthrough is presented in four steps in Figure 3 (with no rule in effect the eight built-in rules and the critic did not question the premise; the user corrects; the rule is compiled and confirmed; with the rule in effect the answer passes grounding against nine rules and the critic accepts after one revision), and the authors note that in this example the correction originated with the user.

The two imaging tabs take a route of lightweight grounding and integration: on the vision side a bone-region head is trained on frozen BiomedCLIP features (MLP 92.6% accuracy, macro-F1 0.918; linear probe 89.6%, 0.886), paired with a nearest-neighbour out-of-distribution guard (threshold at the first percentile of MURA validation scores, 0.827) and image-embedding correction memory (storing the original plus 90/180/270-degree rotations and a horizontal flip, matched at a threshold of about 0.9); on the mechanics side the earlier D2IM model is integrated to predict displacement and strain fields from a single undeformed micro-CT slice with no finite-element model or digital volume correlation at inference. Against the foundation-encoder-and-instruction-tuning route of PathChat, this work trades some peak capability for negligible training cost, edge execution, and safe degradation beyond the training distribution; against isolated deep-learning mechanics surrogates, it places displacement and strain prediction inside the same web application as retrieval and reasoning. Vision metrics are reported on MURA's held-out validation split, which is patient-disjoint from training, with every region individually exceeding 85%; the out-of-distribution guard separates in-scope radiographs (nearest-neighbour cosine 0.83 to 0.99, median 0.95) from out-of-scope images (0.47 to 0.50); the abnormality head reaches only 74.2% accuracy with 0.61 recall on abnormal studies and is explicitly labelled preliminary.

Perspective

The scope is explicitly bounded by the authors: the full-text corpus is restricted to open-access PDFs reachable through OpenAlex and the resolution pipeline, with paywalled articles present in metadata but absent from the body-text index; the knowledge graph serves only as a one-hop evidence channel for the critic and never as an authoritative reasoner; the vision classifier is valid only within upper-limb radiography (MURA), and the out-of-distribution guard withholds predictions on clinical CT, MRI, and micro-CT, which fall back to the bare VLM; the mechanics model was trained on a particular vertebral micro-CT dataset and generalises most reliably to comparable slices. The system is positioned as a research and educational tool for human-in-the-loop use, not a diagnostic device. The architecture also enables next steps: the authors plan to broaden the vision classifier to lower-limb radiographs and clinical CT datasets behind a modality router, add PathChat-style instruction fine-tuning of the VLM with LoRA, and add cross-modal retrieval, automatic context compression, a D2IM retraining command-line interface, and digital volume correlation engines alongside the D2IM prediction in the Mechanics tab.

Open questions the authors identify include: the knowledge graph, although cleaned, is machine-extracted, has not been verified edge by edge for the fracture domain, and remains fragmented; the grounded-reasoning benchmark comprises only fifty questions run once per condition, so significant differences rest on small margins; the link between prompt structure and evidence use is provisional because the underlying differences were not significant; questions were generated by a language model from corpus passages and checked only for the source quantity rather than by domain experts, and one question states its quantities in its own wording and so tests calculation alone; the keyword retrieval conditions were tuned on the same questions and are not yet part of the deployed system, so a held-out question set is needed to confirm them. In addition, the out-of-distribution and correction-memory thresholds must be set conservatively, and the abnormality head must not be interpreted as a clinical flag. For this reading, the body text, tables, and appendix prompts were loaded, but Figures 1 through 10 are image content whose internal detail cannot be verified from the text, so quantitative descriptions tied to those figures rest on what the body text and table notes state.

Sources