Skip to main content
Back to timeline
arXivSource publication:

AutoDecompiler turns binary decompilation into feedback-driven multi-turn refinement with reinforcement learning, consistently beating single-turn baselines at matched scale and input while improving behavioral re-executability

Synopsis

The work presents AutoDecompiler, a decompilation-specialized LLM trained with reinforcement learning that reframes binary decompilation from single-turn generation into an iterative refinement process driven by compilation, execution, and input/output testing feedback, with decompilation-specific rewards covering code validity, recompilability, execution consistency, and semantic fidelity, plus stage-aware diagnostic feedback from compiler errors, execution failures, and failed test cases and progress-aware trajectory rewarding with turn-aware advantage reweighting; across input settings, model scales, and benchmarks it consistently outperforms its single-turn counterparts at the same model size and input setting, with clear gains in behavioral re-executability.

Source-provided article image: Binary Decompilation LLM with Feedback-Driven Multi-Turn Refinement
Fig. 1 ·

Fig. 1: Overview of feedback-driven multi-turn binary decompilation.

arXiv

Interpretation

Binary decompilation is reformulated as feedback-driven multi-turn iterative refinement rather than one-shot code generation. Most prior LLM-based decompilation methods follow a single-turn generation paradigm: given assembly code or decompiler-produced pseudo-code, the model generates one output and stops; here the model revises generated code based on compilation, execution, and input/output testing feedback. The abstract explicitly contrasts the single-turn generation paradigm with the paper's iterative refinement process, noting that single-turn output may appear readable or even compile successfully yet still deviate from the original binary's behavior and mislead downstream analysis.

Decompilation-specific rewards are designed to capture code validity, recompilability, execution consistency, and semantic fidelity. The reward signal is extended from generic code-generation quality to decompilation-specific dimensions of compilability, executability, and behavioral consistency, aligning the reinforcement learning objective with functional correctness in decompilation. The abstract lists the four reward dimensions and states they were designed to enable this iterative process.

Stage-aware diagnostic feedback is constructed, together with progress-aware trajectory rewarding and turn-aware advantage reweighting. Feedback comes from compiler errors, execution failures, and failed test cases and is organized by stage; trajectory-level reward tracks progress and advantage reweighting adjusts by turn, to encourage beneficial revisions while suppressing regressions. The abstract states the feedback sources and the two training mechanisms, and gives their purpose: encouraging beneficial revisions while suppressing regressions.

The AutoDecompiler family is trained and evaluated across settings, consistently outperforming single-turn counterparts at the same model size and input setting, with clear improvements in behavioral re-executability. Relative to single-turn baselines, the improvement is on behavioral re-executability as a functional-correctness measure, not merely code readability or compilability. The abstract reports evaluation across different input settings, model scales, and benchmarks, and states consistent outperformance of single-turn counterparts under the same model size and input setting.

Perspective

The work targets the concrete task of binary decompilation and applies to settings where compilation, execution, and input/output testing feedback can be obtained, such as vulnerability discovery, malware inspection, and executable-only program understanding. Its reported benefit is framed as a comparison at the same model size and input setting, namely improved behavioral re-executability relative to single-turn counterparts. The method's value lies in bringing program feedback into the reinforcement learning training loop, so its precondition is that feedback signals can be constructed and executed.

The abstract does not give specific benchmark names, model-size numbers, sample sizes, effect magnitudes, or statistical tests, so the absolute size of the improvement and its stability across binary types still require the paper's experimental details. Feedback depends on compilation and execution, so how much multi-turn refinement helps for targets that are hard to execute or for which input/output tests are hard to construct remains an open question. The concrete forms of progress-aware trajectory rewarding and turn-aware advantage reweighting, and their individual contributions, also require the paper's ablations to judge.

Sources