LAM resource-bounded abstraction formalizes LLM agent harness cost and validates communication and reliability predictions on GPT-6 Astra
Related research and updatesSynopsis
The authors introduce the Language Model Agent Machine (LAM), a resource-bounded abstraction that fixes the underlying semantic model while explicitly charging harness-level resources, yielding four classes of results on communication, access, recomputation, and reliability, and testing communication and reliability predictions in controlled and held-out experiments on GPT-6 Astra, including checkpoint optima, policy selection under programmatic checking, and tradeoffs among call granularity, logical input traffic, and reliability on chained MATH tasks.
Figure 1. Overview of the lam: (a) an agent with active context, persistent memory, tools and verifiers; (b) the lam components; (c) the charged quantities (calls 𝑁, token traffic 𝑇in, 𝑇out, memory operations, capacities, target failure probability 𝛿); (d) red–blue pebbling (context blocks red, memory blocks blue, transfers communication).
arXiv · Page 1Interpretation
Introduces the LAM resource-bounded abstraction that explicitly charges harness-level resources (bounded context, persistent memory, tools, verification, repeated execution) while fixing the underlying semantic model. Existing notions of model capability do not quantify the computational resources these mechanisms consume; LAM brings harness overhead into an analyzable framework. At the summary level, the abstraction and the structure of four result classes are stated without formal details.
Communication result: LAM execution is instancewise equivalent to red-blue pebbling under simultaneous call-transfer budgets, transferring classical I/O lower bounds to context-memory traffic. Establishes an equivalence between agent execution cost and classical pebbling/I/O lower-bound theory. The summary states the equivalence and the transfer of lower bounds without proof details.
Access and recomputation results: memory interfaces induce asymptotic separations, including a Θ(n) gap between random and non-speculative sequential access on pointer chasing; bit-reversal DAGs require Θ(n²/(C+S)+n) model calls with context capacity C and persistent-memory capacity S. Quantifies when stored intermediate state avoids repeated semantic computation and gives asymptotic separations for access patterns. The summary gives asymptotic expressions without constants or experimental scale.
Reliability results: derives tight stage-local sampling bounds, exact imperfect-verification costs, and a Young-Daly-type checkpoint law with a closed-form optimal verification interval; controlled and held-out experiments on GPT-6 Astra test communication and reliability predictions, including checkpoint optima, policy selection under programmatic checking, and tradeoffs among call granularity, logical input traffic, and reliability on chained MATH tasks. Brings checkpoint and verification-interval optimization into the same resource theory and provides experimental testing. The summary states that experiments cover communication and reliability predictions and several tradeoffs without reporting specific values, sample sizes, or effect sizes.
Perspective
The work targets LLM agent harnesses characterized by bounded context, persistent memory, tools, verification, and repeated execution, and applies to settings that trade off call granularity, logical input traffic, and reliability, such as chained MATH tasks and policy selection under programmatic checking. Its conclusions assume a fixed underlying semantic model, so they apply to analyzing harness-level resources rather than model capability itself.
The summary does not provide proof details, experimental scale, specific values, or effect sizes, so the constant factors of the asymptotic bounds and the experimental coverage cannot be judged from the summary. The applicability conditions of the closed-form checkpoint law and optimal verification interval in real harnesses, and whether the tradeoff conclusions on chained MATH tasks generalize to other task types, remain open questions that require reading the original text.
