Skip to main content
Back to timeline
arXivSource publication:

JEVDB pairs typed decision models with Semantic Bloom Filters to reach the lowest latency on all 21 SemBench queries

Related research and updates

Synopsis

JEVDB is a scalable semantic database system that uses fast, typed decision models for semantic filters, joins, classification, and ranking while selectively escalating uncertain cases to generative LLMs; it combines Yannakakis-style semijoin reduction with Semantic Bloom Filters (SBFs) to cut semantic-join work, achieving the lowest latency on all 21 evaluated SemBench queries and the lowest cost on 19 with competitive answer quality, and on the TPC-DS-derived Shelob workload with joins scaling to 540K candidate pairs it completes all queries with 95.7%-97.5% mean F1, where SBF screening removes 87.4% of candidate pairs before semantic evaluation and reusable condition-index scoring further reduces reasoning-model escalations by 55.2%.

Source-provided article image: Prune First, Decide Fast: Scalable Semantic Query Processing with JEVDB
Figure 1 ·

Figure 1 : JEVDB architecture and query execution workflow. Semantic SQL queries are compiled into relational plans and typed decision primitives. Tier 1 reduces candidates through relational predicate pushdown, Semantic Bloom Filters, relational Yannakakis semijoin reduction, and input deduplication. Tier 2 evaluates surviving inputs with the Jev decision engine, while uncertain decisions are escalated to a frontier LLM in Tier 3.

arXiv

Interpretation

JEVDB assigns discrete relational decisions such as semantic filters, joins, classification, and ranking to fast, typed decision models, escalating only uncertain cases to generative LLMs. Relative to current semantic database engines that rely heavily on autoregressive LLMs for discrete relational decisions, this design moves high-latency, high-cost generative inference off the common path onto an exception path. A system-design and positioning statement at the abstract level; details of each decision model and of the escalation threshold are not given.

To reduce semantic-join work, JEVDB combines exact Yannakakis-style semijoin reduction over relational structure with Semantic Bloom Filters, which use registered necessary conditions to screen candidates across latent semantic edges. It pairs semijoin reduction from classical relational algebra with necessary-condition screening over latent semantic edges, forming a pre-join pruning mechanism. The abstract gives the mechanism and a quantified effect: SBF screening removes 87.4% of candidate pairs before semantic evaluation.

On SemBench, JEVDB-Flash achieves the lowest latency on all 21 evaluated queries and the lowest cost on 19, while maintaining competitive answer quality. Relative to existing engines, it reports leading results across both latency and cost over a broad query set. The abstract reports latency comparisons across 21 queries and cost comparisons across 19, and describes answer quality as competitive without giving specific quality values.

On the TPC-DS-derived Shelob semantic-join workload, where joins scale to 540K candidate pairs, JEVDB completes all queries with 95.7%-97.5% mean F1, and reusable condition-index scoring reduces reasoning-model escalations by 55.2%. It extends evaluation from SemBench to a larger-scale semantic-join workload and quantifies the gains from pruning and escalation reduction. The abstract reports candidate-pair scale, the mean F1 range, and the escalation-reduction percentage; per-query distributions or variance are not reported.

Perspective

The work targets semantic database settings that extend SQL over unstructured data, for query processing that needs semantic filters, semantic joins, classification, and ranking, especially workloads with large join candidate sets; the reported evaluation covers SemBench and the TPC-DS-derived Shelob, and an interactive query simulator, source code, and benchmarks are mentioned as available, supporting reproduction and comparison under similar settings.

The abstract does not give specific answer-quality values, per-query result distributions, escalation thresholds, or how SBF necessary conditions are registered, and these details affect judgment of the method's applicable scope; moreover, Shelob is a TPC-DS-derived semantic-join workload, so behavior under other data distributions and query shapes remains to be observed.

Sources