LLM multi-agent control framework reaches a 93% mean solve rate in a six-module simulated factory, with the orchestrator autonomously rerouting around a silent conveyor-belt fault in all ten runs
Synopsis
The work pairs each factory module with a dedicated LLM-based agent and an MCP tool server that exposes the module's skills via OPC UA method calls, with agents coordinating over MQTT and grounded by real-time factory-state updates; in a simulation of a six-module hexagonal factory it compares orchestrator, peer-to-peer, and monolithic architectures across nine production challenges of increasing complexity, finding that monolithic and peer-to-peer both achieve the highest mean solve rate (93%) while the orchestrator uniquely resolves a silent conveyor-belt fault in all ten runs by autonomously rerouting plates around the blocked segment, and that all architectures exhibit emergent fault-diagnosis behavior without explicit failure-handling logic.
Fig. 1: For each factory module, OPC UA method calls expose the module’s skills, which can be invoked by the module’s corresponding LLM-agent over MCP tools. The agents communicate over MQTT.
arXivInterpretation
Each factory module is paired with a dedicated LLM-based agent and an MCP tool server that exposes the module's skills via OPC UA method calls, with agents coordinating over MQTT and grounded by real-time updates of the factory state. Relative to static programs and centralized programming, the scheme standardizes skill exposure, inter-agent communication, and state injection into a reusable interface layer, allowing offline generation of deterministic production sequences and online handling of runtime faults to coexist. The abstract describes the architecture and communication protocols but does not give interface implementation details or deployment scale.
In a simulation of a physical six-module hexagonal factory, three agent architectures (orchestrator, peer-to-peer, and monolithic) are compared across nine production challenges of increasing complexity. Architecture choice is treated as a controlled comparison variable across a challenge gradient from routine production to silent hardware fault detection, rather than reporting a single architecture's success cases. The abstract reports the comparison design and the number of challenges but not per-challenge scores or statistical tests.
The monolithic and peer-to-peer architectures both achieve the highest mean solve rate (93%), while the orchestrator uniquely resolves a silent conveyor-belt fault in all ten runs by autonomously rerouting plates around the blocked segment. This indicates architectures can tie on mean solve rate yet differ on a specific silent-fault scenario, suggesting architecture choice should match fault type. The abstract gives the 93% mean solve rate and the all-ten-runs resolution record but does not state sample size beyond the run count or variance.
All architectures exhibit emergent fault-diagnosis behavior without any explicit failure-handling logic. This suggests fault-diagnosis capability can arise from the combination of LLM agents, tooling, and state grounding rather than from pre-coded failure-handling rules. The abstract states this observation qualitatively, without quantitative metrics or ablation experiments for the diagnostic behavior.
Perspective
The result targets flexible and reconfigurable automation under small-lot, highly customized production, and applies to factory settings premised on module skills exposed via OPC UA, agents coordinating over MQTT, and real-time injection of factory state. Its value lies in providing a unified module-level agent framework for offline deterministic sequence generation and online runtime fault handling, and in offering a reference for architecture selection: monolithic or peer-to-peer for mean solve rate, orchestrator for silent hardware faults. Standardized MCP tooling, MQTT communication, and state injection are presented as a reproducible foundation for further validation on real production lines and broader module topologies.
What is available is abstract-level information without figures or body text, so per-challenge outcomes across the nine production challenges, the variance around the 93% mean solve rate, behavior on fault types other than the silent conveyor-belt fault, and the criteria for judging emergent fault diagnosis cannot be confirmed from the present text. The gap between simulation and real production lines, the relative performance of orchestrator and peer-to-peer architectures as module count grows, and the adaptation cost of MCP tool servers on heterogeneous equipment remain open questions worth watching.
