Skip to main content
Back to timeline
arXivSource publication:

SynLat aligns compression boundaries to syntax and retains 94.0% of uncompressed accuracy at roughly 44% compression

Synopsis

SynLat reframes chain-of-thought compression as a representation-fidelity allocation problem: constituency syntax partitions a trace into ordered, non-overlapping Syntax-Aligned Units (SAUs), an answer-conditioned Teacher builds progressive KEEP/LATENT targets for a single compression-conditioned Student, and latent units map to a shared discrete codebook, so at inference the Student generates mixed text-and-latent reasoning from only the question and requested compression level. Across two Qwen3 Student scales, Standard- and Long-CoT groups, and three compression levels, SynLat matches or exceeds the strongest evaluated baseline in all 12 task-group aggregates and strictly leads in 11, with overall gains of 3.6/2.6 points at MEDIUM and 7.0/5.5 points at HIGH.

Source-provided article image: SYNLAT: Syntax-Aligned Text-Latent Compression for Chain-of-Thought Reasoning
Figure 1 ·

Figure 1: Compression-decision boundaries for the same reasoning sentence. Token and fixed-window units can fragment local relations, whereas step-level units can bind distinct relations; Syntax-Aligned Units preserve coherent relations while keeping them separately assignable to text or latent form.

arXiv

Interpretation

Formulates budgeted CoT compression as a representation-fidelity allocation problem: deciding which content stays textual and which becomes latent, rather than only deleting or rewriting visible tokens. Prior explicit compression (TokenSkip token pruning, R1-Compress chunk compression, step-entropy filtering, formula preservation) stays within visible text; purely latent methods (COCONUT, CODI) replace the trace entirely and must carry complete derivations without textual anchors, while SIM-CoT reports that scaling implicit reasoning can homogenize latent states. SynLat places the decision on the text-versus-latent allocation itself. Supported jointly by problem formulation, method design, and controlled experiments, as a combination of conceptual reframing and empirical validation.

Introduces Syntax-Aligned Units (SAUs) as the compression decision unit: nested, overlapping constituency spans are flattened into ordered, non-overlapping units that fully reconstruct the trace, so local relations stay intact yet remain separately assignable to text or latent form. Token and fixed-window boundaries fragment phrases, formulas, and local derivations, while step-level boundaries bind content requiring different compression actions; SAUs provide trace-specific units. At roughly 44% achieved CR, Table 2 shows SAUs retain 94.0% of NONE accuracy, with fixed 5-token windows trailing by 2.6 points, guarded windows still trailing by 1.8 points, and step units trailing by 4.4 points. Ablation is a dedicated Qwen3-8B study macro-averaged over five benchmarks, reporting unit average length plus constituency and formula split rates; SAUs show the lowest formula split rate at 5.9 with average length 3.6.

Designs answer-conditioned progressive KEEP/LATENT planning with plan-then-generate two-stage transfer, so one compression-conditioned Student generates mixed CoT at inference from only the question and compression level. The Teacher judges which units are recoverable given remaining textual anchors and context rather than local importance alone, and conditions share a progressive replaceability order so higher compression adds LATENT units without restoring earlier ones. On fixed SAUs, answer-conditioned planning retains 94.0%/89.3% of NONE accuracy at MEDIUM/HIGH versus 92.0%/86.2% for the strongest answer-free alternatives; removing action-plan supervision lowers mean accuracy across the three compressed levels from 73.1 to 70.6. Planning and training ablations are dedicated Qwen3-8B studies macro-averaged over all five benchmarks with reported comparison values.

Under matched output budgets, SynLat matches or exceeds the strongest evaluated baseline in all 12 task-group aggregates and strictly leads in 11, with advantages growing under stronger compression. Overall margins are 3.6/2.6 points at MEDIUM for Qwen3-8B/14B and widen to 7.0/5.5 points at HIGH; for Qwen3-8B at HIGH the margin is 4.7 points on Standard-CoT and 8.5 points on Long-CoT. Baselines outperform SynLat in a few settings: MATH-500 for Qwen3-8B at LOW, and ProofWriter at LOW and StrategyQA at MEDIUM for Qwen3-14B. Main results report accuracy and achieved CR across two Student scales, five benchmarks, and three compression levels, averaged over independent training and generation seeds, with data, computation, and evaluation held fixed.

Perspective

The result targets settings that must deploy long CoT reasoning under constrained output budgets, in languages and tasks where constituency parsing is available and Teacher targets can be built offline. Validation covers two Student scales (Qwen3-8B and Qwen3-14B) on MATH-500, StrategyQA, OlympiadBench, WorldTree V2, and ProofWriter; LOW/MEDIUM/HIGH target 25%/50%/75% SAU replacement, with overall achieved CRs of about 23%/44%/65%. For readers reusing the idea, the transferable parts are the design principles of syntax-aligned allocation units and answer-conditioned progressive planning with two-stage transfer; latent units use a shared finite codebook built from frozen SpanBERT encodings followed by IncrementalPCA, normalization, and spherical mini-batch k-means, and the Student receives special tokens with trainable embeddings rather than offline span vectors.

A careful reader would still watch: SAU construction depends on constituency parse quality, and parser behavior on formula-dense or non-English traces is not elaborated; the codebook is fitted on a task-independent corpus, and sensitivity of latent-unit recoverability to codebook size and coverage is not analyzed; LOW/MEDIUM/HIGH specify unit replacement ratios rather than output CR, and task-level achieved CR varies, so cross-task comparison should follow the paper's achieved-CR selection protocol; main results average independent training and generation seeds, but seed variance is not reported in the body; and the Teacher is a frozen Qwen3-235B checkpoint, leaving the influence of Teacher capability on target quality an open question.

Sources