Skip to main content
Back to timeline
arXivSource publication:

Stratified Consistency Distillation: generating K logical translations with a frontier LLM, stratifying pseudo-labels by entropy, and fine-tuning a smaller model yields consistent gains in Pass@K and Equivalent Logical Similarity

Related research and updates

Synopsis

The work proposes a fine-tuning-based Stratified Consistency Distillation approach that generates K logical translations per input with a frontier LLM and clusters them by semantic equivalence, then selects pseudo-labels by entropy level using majority voting (low entropy), LLM-as-a-Judge (medium entropy), or unification/abstention (high entropy) to fine-tune a smaller model, with experiments showing significant and consistent improvements in both Pass@K and a newly proposed Equivalent Logical Similarity metric.

Source-provided article image: Stratified Consistency Distillation for Natural Language Formalization
Figure 1 ·

Figure 1: Synthetic dataset generation pipeline.

arXiv

Interpretation

It introduces a stratified consistency distillation pipeline that shifts natural-language-to-logical-formula translation from prompt engineering toward a fine-tuning paradigm: a frontier LLM samples K logical translations for the same input, which are then clustered by semantic equivalence. Prior methods predominantly rely on prompt engineering, which is difficult to scale across domains and input formats; this work draws on the success of fine-tuning in other model adaptation and alignment applications, turning consistency signals into trainable pseudo-labels. At the abstract level the paper describes a complete three-step pipeline (generate K translations and cluster, select by entropy stratum, fine-tune a smaller model) and reports improvements in both Pass@K and Equivalent Logical Similarity.

It designs an entropy-stratified pseudo-label selection strategy: majority voting for low entropy, LLM-as-a-Judge for medium entropy, and unification or abstention for high entropy. It uses the degree of agreement among model outputs as the selection criterion, routing samples of different uncertainty levels through different processing paths rather than applying a single aggregation rule to all samples. The abstract explicitly lists the three entropy levels and their three corresponding treatments, which is a method-level design statement; specific thresholds and per-stratum sample proportions are not given in the abstract.

It proposes a new Equivalent Logical Similarity metric, used alongside Pass@K to evaluate logical translation quality. Beyond the existing Pass@K, it adds a measure targeting logical equivalence, so evaluation is not limited to whether the final solve succeeds. The abstract states the metric is newly proposed by the authors and reports significant and consistent improvements on both metrics; the metric's definition and computation are not expanded in the abstract.

Experiments show the fine-tuned smaller model achieves significant and consistent improvements in both Pass@K and Equivalent Logical Similarity, indicating that consistency distillation can advance logical translation. It locates the source of improvement in the pseudo-label supervision obtained through distillation rather than in inference-time prompt design, pointing toward a path that can scale across domains and input formats. The abstract summarizes results as "significant and consistent improvements" without giving specific numbers, dataset names, or baseline comparison details.

Perspective

The work targets neurosymbolic reasoning settings that require converting natural language into logical formulas, and it is suited to researchers and engineering teams who want to replace per-task prompt engineering with a fine-tuned smaller model; its setup is to first generate K translations with a frontier LLM and cluster them, then select pseudo-labels by entropy stratum, and finally fine-tune a smaller model, so the benefit presupposes access to a frontier LLM for batch sampling. The reported improvements are on the two metrics of Pass@K and Equivalent Logical Similarity, in the evaluation context of logical translation quality.

The abstract does not give the value of K, the specific entropy thresholds, the datasets and baselines used, or the concrete numbers for Pass@K and Equivalent Logical Similarity, so the magnitude of improvement and the conditions of applicability still require the main text; the definition and computation of Equivalent Logical Similarity also need to be confirmed in the main text. In addition, how "unification/abstention" in the high-entropy case affects the final training sample distribution, and how far pseudo-label errors propagate during fine-tuning, are questions a reader can continue to watch.

Sources