FLINT combines a lookup agent, template retrieval, and foreign-key-chain pruning to lift Text-to-SQL on production financial schemas from below 50% to outperforming multiple baselines with the same LLM
Related research and updatesSynopsis
The work presents FLINT, a domain-specialized Text-to-SQL system for production financial databases that resolves natural-language concepts into question-specific reference table constraints via a lookup agent, retrieves structurally similar query templates from a compact expert-authored bank using embedding-based retrieval, and prunes a large table schema by traversing foreign-key chains; evaluated on two datasets totaling 359 questions over production financial schemas, it outperforms various state-of-the-art baselines using the same LLM and is deployed in production as part of a financial data retrieval service.
Figure 3: Accuracy vs. generation time on ScienceBenchmark (CORDIS, OncoMX, SDSS), under one harness and LLM. Top-left is best—faster and more accurate. FLINT is the most accurate on every database while generating SQL 5 × \times faster than the multi-agent systems: CHESS issues k = 10 k{=}10 candidate generations per question yet trails FLINT’s single pass, and ReFoRCE’s iterative refinement cannot recover semantically wrong filters that still execute. Generation time excludes database execution.
arXivInterpretation
FLINT targets production financial databases where concepts are stored as opaque integer keys, even simple queries require multiple joins, and filter predicates reference opaque IDs, building three cooperating components: a lookup agent, embedding-based template retrieval, and schema pruning by foreign-key-chain traversal. General-purpose Text-to-SQL systems perform strongly on academic benchmarks like Spider and BIRD, where schemas are relatively shallow and column values are often human readable, but fall below 50% on production financial databases; FLINT closes this gap with domain-specialized design. The abstract specifies the functional role of each of the three components and states that schema linking prunes by traversing foreign-key chains rather than relying on name similarity alone; evaluation covers two datasets totaling 359 questions over production financial schemas.
Under the same LLM, FLINT outperforms various state-of-the-art baselines. The comparison holds the LLM constant, indicating the performance difference comes from system design rather than differences in the underlying model. The abstract states that 'FLINT outperforms various state-of-the-art baselines using the same LLM' and describes the evaluation as two datasets totaling 359 questions over production financial schemas.
FLINT is deployed in production as part of a financial data retrieval service. The move from academic benchmarks to production deployment is what distinguishes this work from purely benchmark-oriented studies. The abstract explicitly states the system is deployed in production as part of a financial data retrieval service; no deployment scale, traffic, or online metrics are given.
Perspective
The work targets a specific setting: production financial databases where concepts are stored as opaque integer keys, even simple queries require multiple joins, and filter predicates reference opaque IDs. Within this setting, FLINT offers a deployable path for financial data retrieval scenarios that must turn natural-language questions into executable queries, and its combination of a lookup agent, expert template bank, and foreign-key-chain pruning suits enterprise databases with large schemas and non-human-readable concepts. Evaluation rests on two datasets totaling 359 questions over production financial schemas, so the conclusions apply to such schemas and query distributions.
The abstract gives no specific accuracy figures for FLINT or the baselines, nor does it state the per-dataset question counts, template bank size, lookup-agent resolution success rate, or schema-pruning recall. The system is deployed in production, but the abstract provides no online traffic, latency, or user-feedback metrics. It also reports no ablation, so the relative contribution of each of the three components remains an open question. Readers assessing transferability to their own schemas would need the dataset composition and error analysis from the full text.
