Interpretable NLI with Graphs of Atomic Propositions: What Does Interpretability Cost?
Synopsis
The study builds a fully graph-based natural language inference pipeline in which premise and hypothesis are decomposed into atomic propositions, converted into ConceptNet triples via constrained decoding, and serialised together with a retrieved ConceptNet subgraph for a fine-tuned Qwen3.5-0.8B-Base classifier that never sees the original text; it reaches 89.7% accuracy on SNLI, only 1.9 points below an identically trained text model, matches published RoBERTa-large on ANLI rounds R2 and R3 (48.0% vs. 48.9% and 44.9% vs. 44.4%) while trailing by 16 points on R1, giving an overall gap of 9 to 14 points against its text counterpart, which the authors call the price of interpretability and attribute to representational rather than data limitations.
Figure 1: The four stages of the pipeline. Both sentences of the pair are decomposed into atomic propositions, each proposition is converted into ConceptNet triples under a constrained JSON schema to form the premise graph P and the hypothesis graph H, a third graph K is retrieved from ConceptNet, and the three graphs are serialised in the order K, P, H for the classifier. No natural language text reaches the classification stage.
arXiv · Page 2Interpretation
A graph-based NLI pipeline whose classifier never processes the input text, featuring a 28-relation ConceptNet extraction vocabulary and a constrained JSON schema. Unlike prior work that feeds text and graphs together, removing the text makes the intermediate structure decisive: an error in the graph is an error in the prediction, and every prediction traces back to a finite list of triplets. The four pipeline stages (atomisation, triplet extraction, knowledge enhancement, classification) are documented in the main text and appendix with prompts, schema and reproduction code; extraction yields an average of 3.6 triplets per proposition.
A ceteris paribus quantification of the price of interpretability using the same backbone, optimiser and nominal epochs. It turns interpretability cost into a measured quantity rather than a qualitative claim: −1.9 points on SNLI and −9 to −14 points on ANLI. The graph model reaches 0.897±0.006 on SNLI over three seeds against 0.916 for the text control; the zero-shot control shows the 0.8B base is at chance on both text and graphs (0.343 on SNLI).
A data-scaling experiment arguing the gap is representational rather than data-limited. Adding roughly 600,000 training pairs from MNLI and FEVER-NLI yields only about +0.6 points on ANLI, indicating that scaling data is not the route to closing the adversarial gap. After two-phase training, ANLI R1–R3 reach 0.568, 0.479 and 0.461, against 0.565, 0.462 and 0.463 without the additional data.
Ablations showing external knowledge is neutral on SNLI while graphs and text are complementary, with text plus graphs reaching 0.921 against 0.916 for text alone. It separates the contribution of retrieved external knowledge from that of modality combination, noting that multi-hop retrieval may be redundant on simple inference while graphs still add value as a complement to text. Both ablations use the same Qwen3.5-0.8B backbone and protocol (seed 0): P+H+K gives 0.892 and P+H gives 0.895.
Perspective
The result applies to English NLI settings that aim for auditable intermediate structures: on SNLI-style single-sentence caption premises the graph pipeline nearly matches the text model, and on adversarial ANLI it is competitive with published RoBERTa-large on R2 and R3. It lets researchers and auditors localise a prediction to specific triplets and feed a corrected line back through the classifier; the authors also note that the interpretability claim still needs direct testing through a human study in which annotators repair a wrong prediction by editing the triplets.
Readers may still watch: the contribution of the ConceptNet subgraph K on ANLI was not ablated, so whether multi-hop retrieval helps adversarial inference remains open; the atomiser degenerates into repetition loops on 7% of ANLI sentences, the cleaning step is itself lossy, and coreference errors are silently frozen into the premise graph; temporal and spatial granularity, comparative and scalar predication, and propositional attitudes and modality have no faithful target in the 28-relation vocabulary; and interpretability is currently qualitative rather than quantified by human readability measures.
