Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics
Synopsis
The work proposes dynamic context adaptation: a validation-generation loop in which a validation agent extracts structured diagnostic feedback from execution traces, a generation agent proposes multiple candidates per iteration, a knowledge graph supplies semantic constraints, and simulated annealing performs non-greedy selection, targeting the LLM code-generation limitation the authors call static binding; across eight problems the method outperforms zero-shot, Reflexion, and OpenEvolve on seven of eight problems at both 300 and 600 evaluations (p < 0.01), achieves the best score at 1000 evaluations on the primary motivating problem of cross-coupled optimization (0.694 vs. 0.681, p = 0.019, d = 0.52), and ablations identify structured execution feedback as the primary driver.
Fig. 2: Left: validation-generation pipeline. Right: SA selection with Metropolis acceptance and geometric cooling.
· Page 6Interpretation
It presents a formal taxonomy of meaning dependencies among software components that characterizes the LLM static-binding limitation: independent, sender related to receiver, receiver related to sender, and cross-correlated sender and receiver, each with a corresponding relation. Prior work often attributes LLM code-generation failures to prompting or search strategy; this work makes explicit the relation in which the meaning of one routine is defined by the runtime behavior of another, and notes that static linking of concepts implicitly assumes mutual independence of meanings. A conceptual and formal contribution, illustrated by the maze-navigation and cross-coupled-optimization examples and the dependency structures in Figure 1, without statistical testing.
It proposes a two-agent, knowledge-graph-based architecture: a validation agent executes the current candidate and emits structured diagnostics (failure modes, missed opportunities, constraint violations), a generation agent proposes k candidates per iteration using the incumbent and KG-derived constraints, and simulated annealing selects among them. Relative to Reflexion's single-trajectory verbal reflection with greedy acceptance, the architecture adds structured diagnostic extraction, multi-candidate generation per iteration, a KG semantic layer, and non-greedy selection; relative to AlphaEvolve-style methods, its stated aim is sample efficiency at low evaluation budgets. Supported by the architecture description and the qualitative comparison in Table 1; experiments use locally served Qwen2.5-72B with SA parameters T0=2.5, alpha=0.85, Tmin=0.01, about 333 iterations x 3 candidates (approximately 1000 evaluations) per problem and 10 random seeds.
It provides empirical evidence of sample efficiency: at 300 and 600 evaluations the method outperforms OpenEvolve on seven of eight problems (all p < 0.01), with the empirical crossover against population-based search located at roughly 600-800 evaluations. It frames the trade-off between feedback-driven early convergence and population-based asymptotic exploration as complementary, and reports the empirical crossover location (about 650 evaluations for circle packing, about 710 for maze navigation). Table 3 and Figure 4 report mean plus or minus standard deviation at 300, 600, and 1000 evaluations over 10 seeds with paired Wilcoxon signed-rank tests; function minimization is the exception (p = 0.284 at 300 evaluations).
Ablations show structured execution feedback is the largest single contributor, followed by SA selection and KG grounding: removing feedback reduces performance by about 28 percent (maze), 31 percent (circle packing), 12 percent (function minimization), and 11 percent (cross-coupled), all p < 0.001 with d > 1.1. It decomposes system gains across feedback, selection strategy, and semantic layer, noting that without diagnostic feedback the generation agent degrades toward near-random candidate generation. Table 4 reports mean plus or minus standard deviation and effect sizes for four conditions at 300 evaluations; SA matters more on high-variance discrete problems (maze minus 16 percent, p < 0.001, d = 0.89) and the KG gives a moderate gain (maze minus 15 percent, p = 0.001, d = 0.78).
Perspective
The results target routine-level code generation whose correctness depends on joint execution behavior across components, especially when the evaluation budget is limited (on the order of hundreds of evaluations) and execution traces and test vectors are available; empirical results cover explicitly coupled problems (maze navigation, cross-coupled optimization) and standard benchmarks with latent execution-dependent coupling (circle packing, TSP, filter design, online judge programming, symbolic regression). The method runs on a locally served open-weight model, with seven of eight problems completing in under 20 minutes, which fits a developer iteration cycle.
The validation agent diagnoses surface-level symptoms rather than performing counterfactual reasoning, and discrete action spaces show high variance because single-token changes produce large behavioral shifts; feedback scope is bounded by the test suite, so latent bugs outside the execution log escape detection; and the single-incumbent design limits diversity on multi-basin landscapes. The crossover location (roughly 600-800 evaluations) and the per-problem numbers come from this paper's experimental setting, so their position may shift with other models, budgets, or problem distributions, which is worth watching in follow-up work on more models and multi-hop dependency chains.
