Skip to main content
Back to timeline
arXivSource publication:

Compact domain footprints in a frozen embedding space enable generative replay for continual pathology report generation without storing slides or patch exemplars, outperforming exemplar-free and limited-buffer rehearsal baselines on multiple public continual learning benchmarks

Synopsis

The work introduces an exemplar-free continual learning framework for whole-slide-image-to-report generation: it builds a compact domain footprint per domain in a frozen patch-embedding space (a k-means codebook, a slide-level code histogram bank, patch-count statistics, and a report-style prototype), uses it to synthesize pseudo-WSIs whose pseudo-reports come from an immediate teacher snapshot for generative replay, and conditions the language model through a style prefix; across multiple public continual learning benchmarks the approach outperforms exemplar-free and limited-buffer rehearsal baselines and supports domain-agnostic inference without explicit domain identifiers.

Source-provided article image: Footprint-Guided Exemplar-Free Continual Histopathology Report Generation
Fig. 1

Fig. 1. Overview of the proposed continual WSI report generation framework.

· Page 3

Interpretation

Represents each domain with a compact domain footprint in a frozen patch-embedding space, enabling rehearsal without storing raw WSIs or patch features. Rehearsal-based continual learning typically retains past samples or features, whereas this work replaces sample storage with a k-means codebook, a slide-level code histogram bank, and patch-count statistics. The method section specifies the footprint composition F(t) = {C(t), H(t), µ(t)_N, σ(t)_N, r(t)}, with the codebook learned by k-means clustering and the histogram computed as normalized code assignments.

Enables footprint-based generative replay: pseudo-WSIs are synthesized by sampling patch counts and histograms, and pseudo-reports from an immediate teacher snapshot supervise the updated model. Unlike replay requiring real past samples, this method uses an immediate teacher to provide pseudo targets without ground-truth reports, and because latent generation operates in the frozen space it avoids generator drift across domains. The paper details the pseudo-WSI synthesis procedure (N ∼ N(µ(j)_N, (σ(j)_N)^2), sampling code indices from the categorical distribution defined by the histogram) and the teacher-based pseudo-report training objective, and reports performance comparable to exemplar replay.

Handles shifts in reporting conventions via per-domain report-style prototypes and style-prefix tokens, and infers the most compatible descriptor directly from the slide signal at inference, enabling domain-agnostic generation. Prompt-based continual learning typically assumes task or domain identity at inference to select the appropriate prompt, whereas this method requires no explicit domain identifier to choose a style descriptor. The method section states the prototype is obtained by mean-pooling token embeddings from a frozen text encoder followed by ℓ2 normalization, mapped to prefix embeddings by a learned linear projection, with style-token positions masked from the loss.

On multiple public continual learning benchmarks, the approach outperforms exemplar-free and limited-buffer rehearsal baselines. It positions footprint-based generative replay as a practical solution for deployment in evolving clinical settings, measuring report quality and retention across domains under domain-incremental scenarios. The abstract and introduction report advantages over baselines, but the loaded text does not provide specific dataset names, metric values, or statistical test details.

Perspective

The framework targets domain-incremental settings: training proceeds over domains (episodes), and at each episode the model has access only to the current dataset Dt and cannot revisit earlier samples. It applies to any WSI-to-report generator that consumes patch embeddings, instantiated here with a HistoGPT architecture. Its design goal is storage efficiency without retaining raw slides or patch exemplars, and it selects a style descriptor at inference without explicit domain identifiers, making it suited to institutional settings where data are governance-restricted and domain identifiers are missing or inconsistent.

The loaded text is the full text, but the experimental section's specific dataset names, evaluation metric values, ablations, and statistical test details are not presented in the text, so the magnitude and stability of the reported advantage over baselines cannot be verified here. In addition, the specific values and sensitivity of hyperparameters such as codebook size K, number of style tokens M, and replay weight λ, as well as the actual effect of Gaussian noise on diversity in pseudo-WSI synthesis, require further confirmation in the original experimental section.

Sources