Skip to main content
Back to timeline
Frontiers in PharmacologySource publication:

Without touching a line of SAS source, a metadata layer turns a 558-macro clinical reporting library into LLM-readable JSON, with 11 of 14 real reports at 80%+ cell-level parity

Synopsis

The study presents a non-destructive metadata-layer framework (bridge map, typed Intermediate Representation, orchestrator) that re-exposes a legacy clinical reporting library's outputs as machine-readable JSON without modifying validated SAS source, validated on a 558-component, 372,698-line industrial SAS macro library: immediate AI readiness under coexistence mode, an optional 92% reduction in proprietary code, cell-level parity of 80% or above on 11 of 14 report types from internal Phase III study PROT008-SR1 (mean 82.7%, best 99.2%), 100% parity across 5 reports and 4,764 cells on the public CDISC CDISCPilot01 benchmark, and LLM experiments covering table summarization, adverse event anomaly detection, and trial configuration generation.

Source-provided article image: A non-destructive methodological framework for modernizing legacy clinical reporting systems for AI-driven pharmacoinformatics: a SAS case study
Figure 1

a coverage matrix. Figure 1 presents the taxonomy breakdown for the case-study library. The methodology requires no proprietary tools and is applicable to any pharmaceutical SAS macro library.

· Page 7

Interpretation

The framework decouples AI readiness from source-level change: a bridge map (365 entries, 1:1 coverage of every callable component) plus a typed IR plus a Python orchestrator wrap the legacy library, whose files are read but never written, and whose outputs are re-exposed as structured JSON consumable by LLMs, R, and Python. Prior options assumed AI readiness required changing legacy source and therefore triggered re-validation cost; this work moves the deployment surface from SAS source to metadata, making AI integration a Day-0 capability rather than the end state of a multi-year rewrite. Implemented on a 558-component, 400-file, 372,698-line industrial library; coexistence mode yields machine-readable output with the legacy component count unchanged, and consolidation is an opt-in incremental upgrade.

The typed IR is the central architectural innovation: ir_cells records per-cell raw numeric cell_value, formatted cell_formatted, and cell_type from a controlled vocabulary (INTEGER, DECIMAL, PVALUE, PERCENTAGE, TEXT, HEADER, LABEL, FOOTNOTE, EMPTY), ir_structure defines row and column dimensions, and cl_ir_reconcile traces cells back to source statistics within a default 1×10⁻¹⁰ tolerance. Legacy libraries had no machine-readable layer between statistical computation and rendered output, and RTF encodes visual formatting rather than semantic content; the IR separates semantic classification and numerics from layout, enabling cell-level QC, format-agnostic rendering, and reconciliation against source compute output. The IR is a byproduct of every compute-to-render handoff; a single IR feeds four render engines (RTF, PDF, HTML, JSON), and render engines contain no statistical logic.

On 14 report types from internal Phase III study PROT008-SR1, 11 cleared the 80% cell-level parity threshold (mean 82.7%, median 89.6%, best 99.2%), and the 3 below threshold stem from the legacy system using PROC TRANSPOSE to pivot treatments into rows, a layout-geometry difference rather than a computational error. The paper supplies a reusable 12-category divergence taxonomy and a seven-gate (Gate A–G) verification workflow, with all 72 fixes applied at the framework layer, producing fix-once/apply-many leverage that raised parity from an initial 8/14 to 11/14. The 14 report types span adverse events, baseline characteristics, disposition, compliance, ECG, laboratory, risk management, and listings; 19 of 365 bridge map entries were validated end-to-end and the remaining 346 through unit-level structural checks at Gates A–D.

On the public CDISC CDISCPilot01 benchmark, 5 report types achieved 100% cell-level parity (4,764 cells, 0 mismatches), and an LLM (Claude Opus 4.6) on the IR completed three tasks: table summarization with all 74 cells numerically correct, adverse event anomaly detection hitting 5 of 5 expected clinical patterns with zero false positives, and generation of valid YAML mapping all five SAP requirements. These AI capabilities are emergent properties of the architectural separation requiring no additional implementation; the two table inputs' IR was produced by legacy SAS macros running unchanged under coexistence mode, showing AI readiness is available on Day 0 of adoption. The public dataset is externally verifiable, with ground truth checked against published CDISCPilot01 summary statistics (e.g., Age Mean Placebo = 75.2 consistent with published 75.21; N = 86/84/84 matching published enrollment); the AI experiments are single-model proof-of-concept without a controlled IR-versus-RTF accuracy comparison.

Perspective

The framework targets pharmaceutical organizations with large validated SAS macro libraries that want AI-readable output without abandoning their existing computational investment, in settings of regulatory-submission TFL generation and pharmacovigilance analysis. Coexistence mode (wrap-only, no source change) is the component most likely to transfer directly, because it depends only on the existence of well-defined macro entry points rather than on the library's internal structure; consolidation is an optional incremental path whose reduction is proportional to the portion an organization elects to consolidate. The IR and JSON export are identical in both modes, so AI workflows can start on Day 0 of adoption. The authors also position the framework as a pragmatic stepping-stone toward CDISC ARS: the IR's report_id, execution_id, cell_type, and ir_structure already correspond at the structural level to ARS result display identifiers, result value types, and row/column specifications, and ARS alignment can be introduced incrementally at the IR layer without disturbing the legacy compute layer.

The parity evidence comes from two complementary tracks, and the authors state explicitly that together they should be read as evidence of mechanism rather than of universal correctness: the real-data track reached 80%+ on 11 of 14 reports and the public track 100% on 5 of 5, but neither exhausts the edge cases encountered across regulatory submissions (missing data patterns, protocol deviations, complex visit structures, organization-specific derivation rules). Only 19 of 365 bridge map entries were validated end-to-end, with the rest at unit-level structural checks. The AI demonstration used a single LLM, performed no controlled IR-versus-RTF accuracy comparison, and included no adversarial testing; the authors list multi-model benchmarking and a formal IR-versus-RTF accuracy study as future work. The framework has not undergone formal IQ/OQ/PQ qualification, though the legacy library stays within its existing validation envelope and only the framework layer requires qualification. The software metrics (LOC, parameter count, nesting depth, coupling, cohesion) are indirect quality measures, and direct measures such as defect rates, mean time to resolution, and programmer productivity were not collected. In addition, the industrial case study's modernized framework source is not publicly available due to organizational policy; what is open-sourced is the IR schema, the YAML specification for 5 public-benchmark report types, and the parity validation harness.

Sources