Skip to main content
Back to timeline
arXivSource publication:

DM³-Nav: Decentralized Multi-Agent Multimodal Multi-Object Semantic Navigation

Synopsis

This paper presents DM³-Nav, a fully decentralized multi-agent semantic navigation system supporting multimodal goal specification (category, language, and image) and multi-object episodes, where robots coordinate solely through ad-hoc pairwise communication exchanging local maps, goal status, and navigation intent without a central coordinator or shared global map; evaluations on HM3DSem scenes (HM3Dv0.2 and GOAT-Bench) show it matches or exceeds centralized and shared-map baselines, and a two-robot team successfully located all 8 multimodal targets in a real-world office environment.

Source-provided article image: DM$^3$-Nav: Decentralized Multi-Agent Multimodal Multi-Object Semantic Navigation
Fig. 1 ·

Fig. 1 : Overview of a multi-agent multimodal semantic navigation episode. Two robots (red and green paths) explore an unseen environment to successfully locate four distinct targets: (A) towel [language], (B) table [image], (C) handbag [language], (D) cabinet [category].

arXiv

Interpretation

A fully decentralized multi-agent semantic navigation architecture that eliminates centralized planning and shared global maps while enabling multimodal goal specifications and multi-object episodes. Prior multi-agent methods (e.g., Co-NavGPT relies on a centralized VLM planner with a shared map; MCoCo-Nav performs decentralized reasoning but still depends on a globally aggregated semantic map) support only single categorical goals; single-agent methods (e.g., GOAT) support multimodal multi-object goals but are limited to one robot. This work is the first to combine multimodal goals, multi-object episodes, and fully decentralized coordination. Achieves the highest success rate of 74.6% on HM3Dv0.2 with two agents (vs. Co-NavGPT 66.1%, MCoCo-Nav 71.6%), and validates on two AgileX Scout Mini robots in a real office environment locating 8 multimodal targets, relying entirely on onboard sensing and computation.

Multi-object Multi-agent SPL (MSPL), a metric that extends SPL to settings with multiple robots, multiple objects, and multiple valid instances per object by measuring efficiency relative to the optimal makespan. Original SPL is designed for single-agent, single-object episodes and cannot handle cases where a goal such as 'chair' may be satisfied by any of several instances; MSPL is based on a multi-agent min-max extension of the Generalized TSP (GTSP) and reduces to standard SPL when n=1 and m=1, making it a strict generalization. The optimal makespan is formulated as a Mixed Integer Linear Program (MILP) solved with Gurobi, and since it depends only on scene geometry and goal locations, it can be computed offline once per episode without affecting runtime performance.

An implicit task allocation mechanism combining intent broadcasting and distance-weighted frontier selection reduces redundant exploration without explicit negotiation. Traditional methods such as auction-based allocation and distributed Hungarian method require coordination rounds, connected communication, or synchronous position knowledge, none of which hold under intermittent ad-hoc connectivity; the distance-ratio heuristic evaluates frontiers independently using only the most recently received neighbor positions. Ablation shows communication improves success rate from 68.6% to 74.6% and SPL from 28.9 to 38.2; a four-robot team achieves a 40% relative improvement in success rate over a single robot (34.5% vs 24.6%) while also completing episodes faster.

Identification and resolution of several limitations in the GOAT navigation pipeline that yield substantial performance improvements even in the single-agent case. Improvements include log-odds accumulation replacing element-wise maximum map updates, raycasting-based explored region marking, three-step goal region selection, local planning refinements, and YOLO-based secondary detection verification. On GOAT-Bench multi-object episodes, with predicted semantics SR improves from 26.3% to 44.0% and SPL from 17.5 to 30.7; with ground-truth semantics SR improves from 56.7% to 85.7% and SPL from 40.3 to 61.6%.

Perspective

This work targets multi-robot multi-object semantic navigation in unknown indoor environments, applicable when robots can engage in intermittent ad-hoc pairwise communication and each robot has onboard sensing and computation. The system is designed to be modular, with the map alignment stage (ORB/RANSAC) and frontier selection strategy both replaceable. Real-world validation was conducted with two AgileX Scout Mini robots in an approximately 200 m² office environment, locating 8 multimodal targets in 14 minutes. Future directions include lightweight perception models for edge deployment, heterogeneous robot teams (e.g., combining UAVs and UGVs), and supporting more complex mission specifications through temporal logics.

MSPL decreases with more agents (0.35 for four robots vs. 0.42 for one), which the authors attribute to the optimal baseline assuming perfect task allocation that becomes harder to approximate as team size grows, but the quantitative boundaries of this explanation remain to be further characterized. Real-world experiments involve only two robots in a single approximately 200 m² office environment; communication reliability, map alignment success rate, and coordination overhead under larger teams and more complex environments remain to be observed. Additionally, the individual contribution of each improvement in Appendix A to overall performance is not separately quantified, so readers interested in the impact of a specific module may need to consult the code or subsequent work.

Sources