An evidence-driven human-agent-robot teaming architecture for Earth-independent anomaly triage, demonstrated end to end with Reachy Mini and MARS
Synopsis
The work presents an evidence-driven human-agent-robot teaming architecture for Earth-independent anomaly triage in deep-space crews: agentic AI acts only as a coordinator over explicit state and bounded, inspectable services, with a triage state manager holding hypotheses, evidence provenance and uncertainty, a crew-facing embodied agent eliciting observations and explaining assessment changes, and a mobile robot acquiring localized evidence; two scenarios, a Mars-transit case anchored on an actual ISS ammonia false alarm and a lunar-base power-interface anomaly, illustrate the architecture, and a hardware-in-the-loop prototype using Reachy Mini and an Innate MARS robot exercises the closed evidence loop end to end.
Fig. 1: Human–Agent–Robot architecture and proof-of-concept components. The Agentic Orchestrator coordinates the crew interface and mobile inspection robot with authoritative triage state and bounded services, while system-governed checks and crew authorization constrain operational actions.
arXivInterpretation
An authority-bounded closed evidence-loop architecture that links an unresolved diagnostic distinction to crew questioning or targeted robot sensing, structured evidence return, hypothesis-ranking revision, and a crew-authorized response, while keeping observations, inferences, and actions explicit. Relative to prior work on tool-using language agents that interleave reasoning with external action and on targeted off-world fault investigation, this work treats the robot as a relocatable sensing resource and closes the full loop from evidence acquisition through hypothesis revision to crew review. Developed through two scenario walkthroughs (Mars-transit ammonia leak, lunar surface power-interface anomaly) and exercised end to end in a hardware-in-the-loop prototype, with an append-only trace recording each state transition, tool request, authority check, robot result, explanation, and crew confirmation.
A bounded tool-use pattern that separates epistemic actions, which gather information, from operational actions, which alter vehicle configuration or commit the crew to a consequential response; operational actions follow a deterministic execution path and require crew authorization except for pre-approved safing. The pattern makes explicit that capability access does not confer operational authority: the orchestrator does not own authoritative state, assign diagnostic confidence, validate its own authority, or declare resolution, leaving safety-critical control with deterministic services and the crew. Implemented at the architecture level through typed interfaces and an authority and safety layer that validates schemas, permissions, operational constraints, and confirmation requirements; in the demonstration, action validation rejected direct crew entry and the robot task was issued as a contract specifying target, sensing objective, permitted modalities, completion and stop conditions, and timeout.
A triage state manager that owns a time-indexed authoritative record containing the initiating event, ranked hypotheses and uncertainty, evidence, operational context, robot and task status, applicable procedure context, and an append-only decision trace, with each evidence item recording source, modality, time, location, quality, provenance, and underlying artifact. The design keeps missing evidence, failed acquisition, and a measured-negative result distinct, and states that dialogue history is not authoritative vehicle state, so delayed ground can review the trace and provide advisory analysis. In the demonstration, the robot's measured-negative NH3 result is distinguished by the evidence service from missing data or a failed acquisition; source images remain primary evidence while schema-constrained visual interpretations are stored as derived observations with model and artifact metadata.
An end-to-end hardware-in-the-loop integration demonstration on an NVIDIA DGX Spark, with an OpenClaw-based orchestrator, Reachy Mini as the crew-facing interface, an authoritative triage state manager with bounded backend services, and Innate MARS reached through a ROS2 robot-tasking service, for the ammonia scenario. The demonstration moves the architecture from paper scenarios to a running integrated prototype, with local models providing dialogue and tool use, speech recognition and synthesis, and visual interpretation, and with inter-component coordination using structured state and typed messages while free-form language appears only at the crew interface. The authors state that the prototype demonstrates end-to-end system integration, not improvement in diagnostic or human-performance outcomes, and that quantitative evaluation remains future work; MARS serves as a morphology-independent surrogate for exercising the same task, status, and evidence interfaces, not for validating free-flight mobility.
Perspective
The architecture targets time-critical anomaly triage for deep-space crews when ground support is delayed or disrupted, and is positioned as decision support under unresolved uncertainty rather than autonomous anomaly detection or emergency management. It applies when an alert has fired, the cause is not yet resolved, and the next step should be a targeted observation rather than an immediate repair attempt; the robot is treated as a relocatable sensing resource for gathering evidence from a suspected hazardous volume or inspecting a physical interface not directly observable through telemetry. During crewed operations, consequential actions require crew authorization; during uncrewed periods, only pre-authorized, reversible actions within verified constraints may execute automatically. Delayed ground may review the trace and provide advisory analysis when communications permit. The authors note that future work will connect validated diagnostic and procedure services, expand multimodal sensing, and measure diagnostic accuracy, workload, trust, and decision quality in habitat-analog studies such as CHAPEA.
The prototype demonstrates end-to-end system integration, not improvement in diagnostic or human-performance outcomes, and quantitative evaluation remains future work. MARS serves as a morphology-independent surrogate for exercising the same task, status, and evidence interfaces, not for validating free-flight mobility; the flight scenario assumes an intravehicular free-flyer. The NH3 reading in the demonstration is simulated and the telemetry is synthetic, and the robot target resolves to a predefined map pose. The authors list future work including connecting validated diagnostic and procedure services, expanding multimodal sensing, and testing missing, delayed, and conflicting evidence; measuring diagnostic accuracy, workload, trust, and decision quality in habitat-analog studies also remains to be done. A careful reader would watch how hypothesis ranking and authority boundaries behave under missing, delayed, or conflicting evidence, and how the architecture transfers beyond the ammonia scenario and the specific hardware combination.
