Org-Agent uses a task dependency graph and constraint-aware execution to extend single-user assistants into organizational agents serving multiple users
Synopsis
The work introduces Org-Agent, a constraint-centric reasoning framework that decomposes a task into atomic subtasks and builds a Task Dependency Graph (TDG), schedules subtasks by topological order, and executes each subtask under constraint rules with evidence-acquisition and memory-management tools; it reaches 44.03% average accuracy on GroupMemBench versus 37.72% for BM25 and 38.79% for the strongest memory system Hindsight, and improves MUSES-Bench averages by 8.27, 4.24, and 5.29 points across three backbones, with ablations supporting both dependency modeling and tool use.
Interpretation
The paper frames organizational agents around two complementary capability dimensions: cross-user interaction and decision-making, and cross-user memory and knowledge use, both governed by organizational constraints across user identity, authority, and access permissions; the attribution and temporal validity of information; and conflict-resolution rules and completion requirements for joint decisions. Prior agent systems are largely designed and evaluated for a single user; this work explicitly models a shared agent serving multiple users and treats constraints as prerequisites of an action and conditions to satisfy during execution, rather than concatenating all users' requests or handling each user independently. This is a conceptual and problem-formulation contribution, illustrated with examples such as product launch scheduling and answering a question about a product's login method, and grounded in two existing benchmarks, MUSES-Bench and GroupMemBench.
Org-Agent constructs a dynamic Task Dependency Graph (TDG) whose nodes are atomic subtasks of three types (inquiry, decision, and response) and whose edges encode dependencies among them; the graph is updated, and reconstructed when needed, as new user inputs change the pending subtasks. Unlike concatenating requests or handling users independently, the TDG makes the cross-user dependencies induced by constraints explicit as a directed acyclic graph and adjusts the remaining plan as new inputs arrive. The paper specifies graph modeling and a three-step construction process (subtask decomposition, dependency identification, graph update), with cycles checked and revised by the LLM; in ablations, removing all edges drops GroupMemBench average from 44.03% to 41.61% and MUSES-Bench from 72.39% to 71.10%.
On top of the TDG, the framework derives an execution order by topological sorting so every node runs after its prerequisites, supports layer-wise parallel execution, and executes each node under constraint rules with evidence-acquisition tools (similarity scoring, conditional filtering, relation traversal) and memory-management tools (read and write of accumulated execution memory). Rather than relying solely on the model's reasoning to handle constraints, the paper places constraint handling in tools and scheduling, letting retrieval filter by metadata such as author, role, project phase, and topic and follow explicit links to trace how discussions and decisions changed. Disabling conditional filtering, relation traversal, and memory management lowers GroupMemBench average to 40.40% and MUSES-Bench to 67.10%; a case study shows that filtering by the asking user's identity avoids the BM25 baseline's error of returning another user's decision due to lexical overlap.
Experiments show gains on both benchmarks: GroupMemBench average accuracy of 44.03%, exceeding BM25 by 6.31 points, text-embedding-3-large by 10.74 points, and the strongest memory system Hindsight by 5.24 points; on MUSES-Bench, GPT-4o-mini averages 72.39% (up 8.27 points), DeepSeek-V4.1-Flash 85.02% (up 4.24), and Qwen3-32B 68.76% (up 5.29). The paper reports that existing agent memory systems do not consistently outperform basic retrieval baselines on cross-user memory and knowledge use, with most falling behind BM25; Org-Agent's gains do not require constructing a memory graph before the query arrives, and its performance degrades more slowly as the user group grows. GroupMemBench contains 745 questions across six query categories, and MUSES-Bench contains 1,183 scenarios across four tasks with group sizes from 2 to 20; answer correctness is assessed by an LLM judge and Instruct by rule-based checkers; a cost analysis shows Org-Agent reaches the highest accuracy with 5.5M tokens, fewer than LightMem's 21.6M and Hindsight's 117.5M.
Perspective
The framework targets an organizational setting where one shared agent serves multiple users, suited to workflows with identity, authority, information attribution, and joint-decision constraints, such as product launch scheduling, cross-user information access, and meeting scheduling. The paper validates it on GroupMemBench (745 questions across Finance, Healthcare, Manufacturing, and Technology and six query categories) and MUSES-Bench (1,183 scenarios with group sizes from 2 to 20), with backbones including GPT-4o-mini, DeepSeek-V4.1-Flash, and Qwen3-32B, plus a Qwen3-8B check in the appendix. For teams aiming to upgrade a single-user assistant into an organizational agent, this constraint-centric design with dependency-ordered scheduling and evidence-acquisition and memory-management tools can serve directly as a starting point; the paper also states that code and relevant resources are available via an anonymous repository with a README providing reproduction guidance.
The reported results come from two benchmarks: cross-user memory and knowledge use on GroupMemBench is judged by an LLM judge for answer correctness, and cross-user interaction and decision-making on MUSES-Bench is measured by rule-based checkers and scenario success rates, so transfer to real organizational deployments under these settings remains to be observed. The user-group-size trend comes from the evaluated range of 2 to 20 users, and behavior at larger scales is an open question. Hyper-parameter sensitivity shows that the balance between lexical and semantic weights in similarity scoring affects results, and suitable values across domains and languages remain to be tested. The cost analysis is based on GPT-4o-mini token consumption on GroupMemBench, and cost structure under other backbones and tasks needs further observation.
