Public articles linked to the same research event.
arXiv The work proposes VALSE, a per-sample, non-contiguous Transformer layer-skipping method in which a lightweight difficulty estimator scores each input from the first few layers and drives per-layer gates, alongside three theoretical results: a closed-form expected-FLOPs formula, strict containment of skip-layer models in the full-layer function space, and a structural duality with Mixture-of-Experts; in a prototype-scale evaluation (12 layers, SST-2) all samples were routed to the minimum depth and the Pearson correlation between difficulty and activated layers was negative, so the core difficulty-adaptive routing hypothesis was not verified.
The work proposes VALSE, a per-sample, non-contiguous Transformer layer-skipping method in which a lightweight difficulty estimator scores each input from the first few layers and drives per-layer gates, alongside three theoretical results: a closed-form expected-FLOPs formula, strict containment of skip-layer models in the full-layer function space, and a structural duality with Mixture-of-Experts; in a prototype-scale evaluation (12 layers, SST-2) all samples were routed to the minimum depth and the Pearson correlation between difficulty and activated layers was negative, so the core difficulty-adaptive routing hypothesis was not verified.
The work proposes VALSE, a per-sample, non-contiguous Transformer layer-skipping method in which a lightweight difficulty estimator scores each input from the first few layers and drives per-layer gates, alongside three theoretical results: a closed-form expected-FLOPs formula, strict containment of skip-layer models in the full-layer function space, and a structural duality with Mixture-of-Experts; in a prototype-scale evaluation (12 layers, SST-2) all samples were routed to the minimum depth and the Pearson correlation between difficulty and activated layers was negative, so the core difficulty-adaptive routing hypothesis was not verified.
The work proposes VALSE, a per-sample, non-contiguous Transformer layer-skipping method in which a lightweight difficulty estimator scores each input from the first few layers and drives per-layer gates, alongside three theoretical results: a closed-form expected-FLOPs formula, strict containment of skip-layer models in the full-layer function space, and a structural duality with Mixture-of-Experts; in a prototype-scale evaluation (12 layers, SST-2) all samples were routed to the minimum depth and the Pearson correlation between difficulty and activated layers was negative, so the core difficulty-adaptive routing hypothesis was not verified.