Does per-frame early exit pay off? A dynamic-depth speech enhancer lands on the same latency-quality frontier as static models on STM32N6, with a 26 microsecond per-frame policy
Synopsis
The work supervises every intermediate depth of one causal speech enhancement model and fine-tunes its output heads so deeper outputs are never worse than shallower ones, yielding a family of static models that are more Pareto-efficient than equivalently-sized counterparts trained from scratch on the same budget (up to 0.11 higher PESQ at equivalent compute, matching the best PESQ at 30% less compute); after int8 quantization on an STM32N6 microcontroller, the dynamic enhancer lies on the same latency-quality frontier as the static models, with the policy running in 26 microseconds per frame on the companion Cortex-M55 and splitting the enhancer into separate NPU graphs adding 2.2% latency overhead.
Figure 3: Pareto front for Float32 deployments on the test set.
arXivInterpretation
A training protocol that supervises every intermediate depth of a single causal model and fine-tunes its output heads so that deeper outputs are never worse than shallower ones. Where depth-varying networks typically require separately or from-scratch trained models, this derives an entire family of static models from one model with a depth-monotonicity guarantee. The abstract states the protocol is used to derive a family of static models and compares them against equivalently-sized counterparts trained from scratch on the same budget; training details, data splits, and ablations are not given in the loaded text.
Static models derived with this protocol are more Pareto-efficient than equivalently-sized counterparts trained from scratch on the same budget: up to 0.11 higher PESQ at equivalent compute, and matching the best PESQ at 30% less compute. Quantifies the value of the training protocol as a position on the compute-quality frontier rather than a single quality number. The abstract gives two concrete figures, 0.11 PESQ and 30% compute savings, against counterparts trained from scratch on the same budget; no confidence intervals or run-to-run variance are reported.
After int8 quantization and latency-quality frontier measurement on an STM32N6 microcontroller, the dynamic enhancer lies on the same frontier as the static models rather than trading quality for dynamic execution. Moves the dynamic-depth discussion from theory or simulation to measured int8 deployment on a real microcontroller, addressing the constraint that most devices can only accelerate static int8 graphs. Measured on VoiceBank-DEMAND with an STM32N6; the abstract does not give specific latency values or the number of operating points on the frontier.
The cost of dynamic execution is small: the policy runs in only 26 microseconds per frame on the companion Cortex-M55, and splitting the enhancer into separate NPU graphs adds 2.2% latency overhead. Directly quantifies the scheduling and graph-splitting cost of the implementation path that orchestrates several static graphs with a policy, showing the extra cost of dynamic depth is negligible. The abstract reports two concrete measurements, 26 microseconds per frame and 2.2%; measurement conditions, frame length, and whether worst cases are included are not stated.
Perspective
The result targets engineering settings that deploy speech enhancement on devices able to accelerate only static int8 graphs, such as hearing aids, headsets, and earbuds. It shows that one causal model can yield a family of static models, with a policy selecting depth per frame, and that the latency-quality frontier can be measured in int8 on an STM32N6. For readers planning dynamic-depth inference on similar microcontrollers, the work offers a reusable training protocol and an order of magnitude for scheduling overhead (26 microseconds per frame for the policy, 2.2% latency for graph splitting), serving as a starting point for scheme selection and budget estimation.
The loaded text is abstract-level and does not give training details, data splits, ablations, latency values, or statistical variance, so the stability of the 0.11 PESQ gain and 30% compute saving across model sizes or policies cannot be judged. The depth-monotonicity guarantee holds after output-head fine-tuning; whether it survives int8 quantization is not stated in the loaded text. The measurement conditions for the 26 microseconds per frame and 2.2% overhead (frame length, worst-case inclusion, policy complexity) are also not given, so readers estimating their own system budgets should keep margin. In addition, the conclusions rest on VoiceBank-DEMAND, leaving behavior under other noise conditions and languages an open question.
