Public articles linked to the same research event.
arXiv The work proposes a fine-tuning-based Stratified Consistency Distillation approach that generates K logical translations per input with a frontier LLM and clusters them by semantic equivalence, then selects pseudo-labels by entropy level using majority voting (low entropy), LLM-as-a-Judge (medium entropy), or unification/abstention (high entropy) to fine-tune a smaller model, with experiments showing significant and consistent improvements in both Pass@K and a newly proposed Equivalent Logical Similarity metric.
The work proposes a fine-tuning-based Stratified Consistency Distillation approach that generates K logical translations per input with a frontier LLM and clusters them by semantic equivalence, then selects pseudo-labels by entropy level using majority voting (low entropy), LLM-as-a-Judge (medium entropy), or unification/abstention (high entropy) to fine-tune a smaller model, with experiments showing significant and consistent improvements in both Pass@K and a newly proposed Equivalent Logical Similarity metric.
The work proposes a fine-tuning-based Stratified Consistency Distillation approach that generates K logical translations per input with a frontier LLM and clusters them by semantic equivalence, then selects pseudo-labels by entropy level using majority voting (low entropy), LLM-as-a-Judge (medium entropy), or unification/abstention (high entropy) to fine-tune a smaller model, with experiments showing significant and consistent improvements in both Pass@K and a newly proposed Equivalent Logical Similarity metric.
The work proposes a fine-tuning-based Stratified Consistency Distillation approach that generates K logical translations per input with a frontier LLM and clusters them by semantic equivalence, then selects pseudo-labels by entropy level using majority voting (low entropy), LLM-as-a-Judge (medium entropy), or unification/abstention (high entropy) to fine-tune a smaller model, with experiments showing significant and consistent improvements in both Pass@K and a newly proposed Equivalent Logical Similarity metric.