D-RAC turns 236 PDFs into 1,748 retrieval-ready chunks with one multimodal pass, cutting chunking cost 77.8%–85.6% below frontier-model agentic chunking
Synopsis
The work presents D-RAC, which deterministically normalizes any enterprise document to PDF, applies a single multimodal LLM pass to convert rendered pages into retrieval-optimized Markdown, and then reuses W-RAC's ID-level chunk planning; on the 236-document, 795-page PDF subset of RAG-Multi-Corpus it converted and chunked the whole corpus in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks, reducing chunking-stage output tokens by 95.7%, cutting chunking cost by 77.8% (GPT-4.1) to 85.6% (Gemini 2.5 Pro), reducing chunking time by 75%, and reaching Recall@6 0.798, MRR 0.690, and NDCG@6 0.801 over 762 queries.
Interpretation
D-RAC extends retrieval-aware chunking from web content to arbitrary enterprise document formats: DOCX, PPTX, XLSX, scans, or native PDF are first deterministically normalized to PDF, a single multimodal LLM pass converts rendered pages into retrieval-optimized Markdown, and the rest of the pipeline follows W-RAC's deterministic parsing and ID-level chunk planning unchanged. Where the earlier W-RAC assumed input with recoverable structure, namely HTML that converts deterministically to Markdown, this work generalizes that premise to PDF and to any renderable format, and uses the multimodal model exactly once, for format conversion rather than chunk generation. The paper specifies a four-stage architecture (normalize and render, convert, parse and section, plan and reconstruct), with rendering as a local deterministic operation at 200 DPI capped at 1,568 pixels per dimension, and conversion using Gemma-3 27B/12B in 5-page batches with 5 parallel workers at temperature 0.1.
Retrieval-aware table normalization rewrites every table row as one self-contained declarative sentence using column headers as context, and explicitly forbids merging distinct rows into disjunctive sentences, so each tabular fact becomes independently embeddable and retrievable. The paper identifies the failure mode of dense-vector retrieval over table cells as the loss of row and column context, and therefore pushes retrieval awareness earlier, into format conversion, rather than leaving it to chunking or embedding to repair. The conversion prompt forbids Markdown table syntax and forbids merging distinct rows (for example, Policy Term=16/20 must not become '16 or 20 years'), and a deterministic post-processing pass converts any escaped table syntax into per-row prose; the paper states this builds on its earlier finding that contextualized tabular prose improves summarization and QA over tables.
A recursive sectioning algorithm splits documents that exceed the planning budget at header boundaries and carries the parent-header chain into each planning call, enabling chunk planning over documents of 500+ pages. Rather than handing a whole document to the model at once, it uses a 60-element budget, prefers the coarsest heading level that fits, descends to finer levels only where needed, falls back to fixed-size splitting for header-free regions, and merges adjacent small sections. In a 503-page stress test, a 5,060-element document was planned in 68.7 seconds across 95 parallel section calls, roughly 5% of its conversion time, with parent-header context costing only a few dozen input tokens.
On the 236-document, 795-page PDF subset of RAG-Multi-Corpus, D-RAC completed the full pipeline with zero conversion and zero chunking errors, exceeded fixed-size chunking over rule-based extraction on retrieval quality, matched agentic chunking, and sharply reduced chunking-stage cost. The paper notes that the agentic reference chunks were produced from the corpus's clean structured sources, whereas D-RAC worked from rendered PDF pages, the hardest input format, yet still matched or exceeded agentic chunking on all seven overall metrics. Conversion totaled 3,758.5 seconds (71.7 minutes wall clock including chunk planning), yielding 1,020,219 characters, 5,584 elements, and 1,748 chunks; over 762 queries Recall@6 rose from 0.717 for fixed-size to 0.798, MRR from 0.602 to 0.690, and NDCG@6 from 0.764 to 0.801; chunking output tokens fell from 270,454 to 11,714 (a 95.7% reduction), cost fell 77.8% (GPT-4.1) and 85.6% (Gemini 2.5 Pro), and chunking time fell from 2,167.5 seconds to 541.8 seconds.
Perspective
The result targets the document ingestion and chunking stage of enterprise knowledge bases and applies to any input format that can be faithfully rendered to PDF, including DOCX, PPTX, XLSX, HTML, scanned images, and native PDF. The evaluation corpus is the PDF subset of RAG-Multi-Corpus, covering product sheets, FAQs, policy and procedure documents, parts catalogs, and service guides across five fictional enterprise organizations in automotive, academia and education, cloud services, enterprise technology, and banking and finance. Because those inputs are natively PDF, the normalization stage is the identity in these experiments, so conversion of other formats is supported by the architecture rather than directly measured here. Retrieval evaluation uses 762 queries with supporting-fact ground truth across four organizations; CloudWay-24 has no annotated queries in the reference set. The cost comparison is priced at published GPT-4.1 and Gemini 2.5 Pro rates, and the time comparison cites agentic chunking processing time measured in the W-RAC experiments. The paper states the design extends naturally to entity-aware chunking, graph-based retrieval, and policy-driven chunk recomposition, and that W-RAC and D-RAC together form a unified ingestion foundation for web content and documents.
Conversion quality depends on the multimodal model reading rendered pages; the paper constrains output with verbatim preservation, a no-merge rule, image suppression, and deterministic post-processing, but the conversion and planning prompts sit in the appendix while the body gives only the rule highlights. Retrieval relevance is judged by a rule requiring at least 60% of a supporting fact's content words to appear in the chunk rather than by human relevance labels; the judge is identical across the three systems, but how that criterion affects the relative ordering of query types is not explored in the body. Boolean queries are the one category where agentic chunking retains an edge, which the paper reports as an observation without a mechanism. The extrapolation of cost savings to a 1M-page enterprise corpus is derived from pricing and token structure rather than measured at that scale. Finally, this evidence bundle is full text, but tables appear as text in the body, so alignment of some numeric columns should be checked against the original tables.
