Skip to main content
Back to timeline
arXivSource publication:

Closed-loop framework uses LLM analysis of policy rollouts to auto-correct structured policies, improving imitation learning by up to 15%

Synopsis

The work proposes a closed-loop framework that logs policy rollouts as semantically meaningful tabular data and prompts an LLM to generate diagnostic analysis code, iteratively identifying and correcting suboptimalities in structured policies without human instruction; on car racing and door opening tasks, it improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance.

Source-provided article image: Iterative Policy Refinement through Semantic Rollout Analysis
Figure 1 ·

Figure 1: Overview of policy refinement. 1. Generate the initial structure by querying an LLM with the task description. 2. Update parameters through imitation learning on the demonstrations. 3. Generate analysis code based on the desired behavior of the policy and its structure topology. 4. Collect tabular trajectories , recording observations, actions, and named latent values during policy rollouts. 5. Execute analyses and revise the structure based on the analysis results. The updated program returns to parameter fitting and another refinement cycle. 6. Reinforcement learning after the structure aligns with the demonstrations to go beyond the demonstration performance.

arXiv

Interpretation

A closed-loop framework is proposed that iteratively refines structured policies through LLM-guided analysis of policy rollouts, automatically correcting suboptimalities in policy structure without human instruction. Existing structure generation methods rely on extensive human input or static domain knowledge encoded in LLMs, which may be inconsistent with expert demonstrations; this work instead uses the policy's own rollouts as the feedback signal. Experiments on car racing and door opening tasks report up to 15% improvement in imitation learning performance over zero-shot LLM-generated structures and 75% less compute to achieve the same reinforcement learning performance.

Logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code makes tabular rollout analysis an effective feedback signal for aligning LLM-generated policy structures with expert demonstrations. It converts rollout logs into a tabular form that an LLM can diagnose, rather than relying only on static domain knowledge or human feedback. The abstract reports experimental results on car racing and door opening tasks showing the method can automatically generate good policy structures.

Perspective

The framework targets imitation learning settings that require structured policies and applies to tasks whose rollouts can be logged as semantically meaningful tabular data, such as car racing and door opening. It enables researchers to automatically and iteratively correct policy structures without human instruction and may reduce the compute needed to reach the same reinforcement learning performance. The results apply to the two task settings described in the abstract; applicability to broader tasks, robotic platforms, and real-world environments remains to be verified.

The abstract does not specify the fields of the tabular rollout data, how the LLM diagnostic code is generated and executed, or the iteration stopping criteria, nor does it report sample sizes, random seeds, or statistical significance. Whether the improvements on the two tasks replicate on other tasks and how the compute reduction is measured remain open questions for readers. Because only the abstract is available, figures and experimental details cannot be checked, and the above should be treated as open questions rather than conclusions.

Sources