Skip to main content
Back to timeline
arXivSource publication:

DBRAG combines a table index with query-relevant row reranking to reach 95.8 Recall@5 on Spider and improve multi-table QA

Synopsis

DBRAG decomposes multi-table QA into two stages: an offline table index retrieves candidate tables, query-relevant rows replace random rows in their summaries for LLM reranking, and a program-aided chain-of-thought reasoner selects tables and executes operations over their full contents; on Spider, GeoQuery, and ATIS it improves table retrieval and multi-table QA metrics over the compared baselines.

Source-provided article image: DBRAG: Multi-Table Retrieval-Augmented Generation for Complex Database Queries
Figure 1 ·

Figure 1: The DBRAG execution workflow. The Table Index retrieves top- K K candidates; the Row Index supplies query-relevant context rows. An LLM reranks the candidates to select top- M M tables. These tables are then passed to a program-aided reasoner, which uses a multi-table CoT prompt to generate the final answer.

arXiv

Interpretation

DBRAG explicitly decomposes multi-table QA into multi-table retrieval and multi-table reasoning, using a two-stage retriever (table-index coarse ranking plus LLM fine-grained reranking) and a multi-table-aware chain-of-thought reasoner with program-aided execution tools. Prior work largely targets single-table settings or assumes the relevant tables are already provided; DBRAG addresses the open-domain setting where tables must be found from an entire database corpus. The paper gives a formal task definition, two stated requirements (R1 retrieval diversity, R2 reasoner context size), and full component and prompt descriptions, making the method reproducible.

On Spider retrieval, DBRAG with relevant row context achieves the best average Recall@5 (95.8), above JAR (Coarse-rank) at 93.0, Coarse-rank at 89.5, DTR at 83.0, and Contriever at 80.6. Ablations show that dropping row information (name and schema only, 83.7) or using random rows (94.9) both reduce performance, indicating that query-relevant rows rather than arbitrary rows drive the gain. Spider has 985 test samples, reported by groups of 1- to 4-table queries; the 4-table group contains only six queries, which the authors note.

The reranking LLM affects retrieval: GPT-3.5 (Turbo-0125) averages 95.8, Llama 3.3 (70B-Instruct) 92.5, and Qwen2.5 (32B-Instruct) 90.8, with the gap widening on more complex queries requiring more tables. This comparison holds the retrieval setup fixed, indicating that fine-grained reranking depends on model reasoning capability rather than on index structure alone. Three models compared on the same Spider dataset under default parameters, stratified by the number of tables a query needs.

In end-to-end multi-table reasoning, DBRAG outperforms MTQA and ReadTable on most metrics across Spider, GeoQ, and ATIS, particularly on Spider where ReadTable fails many queries because tables exceed the LLM context limit. DBRAG uses compact summaries to guide table selection and operation construction while execution tools still access the full retrieved tables, so context-row selection does not restrict answer computation. Evaluation uses Table EM, Row EM, Column EM, and Cell EM, with macro-averaging for the latter three to reduce bias from large-table queries; ablations show that removing the selection step or replacing relevant rows lowers answer accuracy.

Perspective

The framework targets multi-table QA where relevant tables must be retrieved from an entire database corpus, and applies to questions whose answers can be obtained through SQL-style operations; the authors state that future work will extend support to heterogeneous data integration such as semi-structured data and knowledge graphs. For practitioners, DBRAG's value lies in concentrating the context budget on compact summaries while execution tools access full table contents, keeping prompt size manageable in databases with many tables and large individual tables.

The computational cost of building and maintaining the Row Index at the scale of billions of records remains open, and the authors leave it, along with dynamic schema updates and production deployment, for future work. Dataset queries involve multiple tables but their answers can be obtained by executing SQL, so open-ended exploration or trend-interpretation queries are not yet evaluated; real databases with noisy, inconsistent, or missing data are also outside the evaluation. The baseline set is limited and does not include more advanced solvers. In retrieval evaluation, the 4-table group contains only six queries, so that subgroup's results warrant caution.

Sources