Skip to main content
Back to timeline
arXivSource publication:

SAP Signavio proposes an "organizational memory" architecture that lifts LLM agents' policy compliance in a procurement process from 30% to 88%–95%

Synopsis

The work introduces the concept of an organizational memory for agentic business process execution—a shared, human-governed, agent-consumable reference layer of organization-specific procedural knowledge—and derives requirements (R1–R9), an architecture covering memory curation and runtime consumption, and an instantiation based on process atoms; in a purchase-to-pay proof-of-concept, the Policy Compliance Rate averaged over 10 scenarios with four runs each reached 88% with GPT-4.1 and 95% with Claude Sonnet 4.5 for the memory-equipped agent, versus 70% and 80% for RAG and 30% for the no-knowledge base setup with both models.

Source-provided article image: Organizational Memory for Agentic Business Process Execution
Fig. 1

Fig. 1: Running purchase-to-pay example: each agent depends on overlapping knowledge that is fragmented across heterogeneous artifacts.

· Page 4

Interpretation

The paper introduces and defines an organizational memory for agentic business process execution: a shared, human-governed, agent-consumable reference layer of organization-specific procedural knowledge about how work should be executed. Prior work either encodes organization-specific knowledge in individual agents' prompts or retrieval setups, or focuses on memory mechanisms for individual agents such as memory tiers, memory streams, and long-term conversational memory; this work treats organization-specific process knowledge as a shared enterprise resource, emphasizing consistent maintenance and governance across agents and processes. Conceptual and architectural argument, grounded in a running purchase-to-pay example (BPMN model, SOP, procurement policy, ERP documentation, knowledge-base article, past execution experience) from which five challenges are derived and nine requirements (R1–R9) are proposed.

The paper proposes an organizational memory architecture with two phases, memory curation and memory consumption, and uses the process atom as its smallest knowledge unit. Unlike unifying sources into plain text or document chunks, a process atom is a self-contained unit capturing exactly one rule, constraint, or responsibility, separating applicability conditions from required action and purpose, and carrying a name, a source reference, and categorized tags grounded in an Enterprise Domain Model. The design choices are mapped to requirements: atomic decomposition supports conflict detection (R3), targeted updates (R7), per-rule traceability (R4), and retrieval of the smallest sufficient set (R5); tag grounding supports context-specific retrieval; curation has a Global Curator performing quality and duplicate/conflict checks, followed by human expert review and approval (R8). The paper states it does not claim or evaluate that a single retrieval mechanism is best.

In the purchase-to-pay proof-of-concept, the organizational memory agent achieved the highest Policy Compliance Rate: 88% with GPT-4.1 and 95% with Claude Sonnet 4.5, above RAG (70% and 80%) and the no-knowledge base setup (30% for both models). RAG helps when the relevant policy is semantically close to the employee request but fails on cross-cutting rules—for example, when an employee requests notebooks without specifying a cost center, the RAG agent creates the PR although the organization requires a cost center for every PR, because that universal rule sits in a finance policy and is not semantically close to the request; organizational memory mitigates this by retrieving task-relevant atoms rather than relying solely on similarity, and also handles scenarios requiring multiple policies to be considered jointly. The evaluation uses 9 manually created PDF documents (5 procurement-related and 4 distractor documents from other functions such as HR) encoding 11 rules, 10 scenarios with four runs per scenario averaged, and rule-based correctness checking against manually created ground truth (Create/Refuse plus material, quantity, and vendor fields); all three setups use the same LLM and system prompt.

The paper reports that the effectiveness of organizational memory depends not only on runtime retrieval but also on the quality of memory curation: the memory agent's remaining errors are mainly over-restrictions, refusing requests that are actually compliant, traced by inspection to an imprecisely extracted atom whose applicability scope is too broad. This shifts the evaluation focus from retrieval strategy alone to the precision of atom extraction and curation, motivating future work on the engineering side (robust extraction mechanisms for knowledge sources beyond PDF documents) and on the management side (responsibility, accountability, lifecycle management, update propagation, and governance at scale). The conclusion comes from inspecting error cases in the proof-of-concept; the paper explicitly cautions that results should be interpreted with caution given the preliminary nature and limited scope of the evaluation, and that organizational memory is not designed to and cannot ensure correct process execution in all cases.

Perspective

The result targets enterprise settings where LLM-based agents execute process tasks, especially steps whose correct decision depends on organization-specific procedural knowledge such as policies, responsibilities, systems, and exception-handling practices; the paper's example is creating purchase requisitions in a purchase-to-pay process, where the agent must decide whether a request complies with organizational rules, create the PR with correct information if it does, and refuse with an explanation if it does not. On the architecture side, the curation phase can take policies, SOPs, process models, system documentation, knowledge-base articles, and traces from previous executions as input, but for simplicity the paper focuses on PDF documents such as policies and SOPs and states that designing extraction services for all possible input sources is beyond its scope. The consumption phase targets runtime agents, where a context request and a Retriever return the smallest sufficient set of atoms, with tag grounding making atoms retrievable by process, activity, role, object type, or organizational context. Human governance is treated as necessary: experts review atom changes and may accept, modify, reject, or resolve conflicts differently, with different departments such as finance, procurement, and compliance owning different rules.

The proof-of-concept covers a single synthetic procurement process and a small policy corpus consisting only of PDF documents; generalization to real enterprise settings with numerous overlapping policies, process models, conflicting sources, and legacy documentation is not yet demonstrated, an open question the paper itself raises. The memory agent's remaining errors are mainly over-restrictions caused by an extracted atom whose applicability scope is too broad, indicating that extraction and curation precision strongly affects outcomes, yet the paper does not provide a systematic assessment of extraction quality. The retrieval mechanism is deliberately left to instantiation, and the paper neither claims nor evaluates that a single mechanism is best, so how reliably the smallest sufficient set avoids both under-retrieval and over-retrieval at enterprise scale remains an open question. In addition, this reading is of the full text, but Figure 2 (architecture overview) and Figure 3 (results chart) are presented as images, so their internal details cannot be checked item by item from the text and the reported numbers follow the body text.

Sources