Nanjing University team releases DSEBench, the first test collection supporting both keyword-plus-example dataset search and field-level explanations
Synopsis
The work formalizes Dataset Search with Examples (DSE) and its explainable extension (ExDSE), builds DSEBench as the first test collection with both dataset-level and field-level human annotations (141 test cases, 7,415 triples, plus 5,699 training cases and 122,585 triples), uses a large language model to generate large-scale training annotations, and adapts and evaluates a range of retrieval, reranking, and explanation methods to establish baselines.
Fig. 1: An example of ExDSE. Left: The user submits a query and a target dataset. Right: The system aims to return positive candidate datasets that ex- hibit both relevance to the query (highlighted in yellow and green) and similarity to the target dataset (highlighted in magenta, cyan, and gray), as opposed to negative ones that are only relevant to the query or only similar to the target dataset. These judgments are supported by specific metadata and content fields.
· Page 2Interpretation
The paper formalizes an information need that combines a keyword query with a few target datasets as examples into the DSE task, and further requires identifying which fields of a candidate dataset indicate relevance or similarity, forming ExDSE. Dataset search had been split into keyword-based retrieval and similarity-based discovery, and existing test collections such as NTCIR, ACORDAR, and WTR target keyword-only input with dataset-level annotations only, so they cannot evaluate a setting that depends on both a query and examples and asks for field-level explanations. The task definition is stated explicitly in the introduction and illustrated in Figure 1 with an education dataset example, where positive candidates must satisfy both query relevance and target similarity while negatives satisfy only one; the field-level requirement is written into the definition as identifying a subset of metadata or content fields.
DSEBench is built on the English subset of NTCIR (46,615 datasets, 92,930 data files), defining each dataset by five fields, title, description, tags, author, and a summary extracted from data files, with content summaries generated to support retrieval and human annotation. Prior test collections did not fold data-file content into a unified field representation; this work detects the real format of each file with python-magic and extracts headers, abstracts, or keys/tags/predicates accordingly, reporting 92% header-extraction accuracy on a random sample of 50 CSV files and 91% on 100 XLSX/XLS samples, with textual abstracts generated by bart-large-cnn. The field definitions, the format distribution table (PDF at 51.02%, and so on), the summary generation pipeline, and the sampled accuracies are all reported in the text; the summary field later shows a selection frequency comparable to title and tags in the explanation evaluation.
Test cases were adapted from 141 highly relevant NTCIR query-dataset pairs, pooled into 7,415 triples to annotate, and judged by 24 students with dataset search experience on a 0/1/2 graded scale for query relevance and target similarity, with indicator fields annotated for relevant or similar datasets. This is the first DSE test collection providing both dataset-level and field-level ground truth; annotation used two independent annotators with a third resolving disagreements by majority vote, yielding Krippendorff's alpha of 0.51 for query relevance and 0.52 for target similarity, above the 0.44 reported on NTCIR. Annotator count, grading scale, arbitration procedure, agreement statistics, and judgment distributions (68.19% irrelevant, 52.47% dissimilar, 9.94% highly relevant, 14.16% highly similar) are given in the text and in Tables 2 and 3 and Figures 3 and 4.
On the training side, GLM-3-Turbo annotated 337,976 triples, and after heuristic filtering 122,585 fully annotated triples from 5,699 training cases were retained, showing training effectiveness close to human annotations in the experiments. Against high annotation cost and scarce training data, the work offers a path of LLM-based training-data expansion with quality control; a manual check of sampled triples reports accuracy rising from 73% and 69% to above 85% for query relevance and target similarity judgments. The quality-control rules, the sampled check size (3,378 triples), and the accuracy changes are reported; in retrieval evaluation the gap between the five-fold and annotator splits is small, which the authors read as LLM annotations being almost as effective as human annotations for training.
Perspective
The test collection targets research and evaluation settings in dataset search: it can be used to train and compare retrieval, reranking, and explanation methods for tasks that must consider a keyword query together with target dataset examples and must give field-level justifications. The authors have released the collection on Zenodo and baseline code on GitHub, and plan to keep adding manually annotated cases, to propose a challenge at ISWC or a shared task at an IR conference, and to extend the methodology to other collections, especially RDF-focused ones. For users, it is suited as a reusable benchmark and a source of weakly supervised training data.
Each query in the test cases is associated with only a single target dataset, so multi-target evaluation settings are not yet covered; pooling simply expands the query by concatenating target dataset fields and does not assess query relevance and target similarity separately; target similarity annotations rely on a general, context-agnostic notion of similarity, whereas criteria in real user contexts may differ. In addition, field-level LLM judgments were not filtered, and the authors describe their error patterns as complex and hard to detect; training-case distributions broadly resemble test cases but come from a different annotation source. These are directions a reader can keep watching when reusing the collection or comparing baseline results.
