Skip to main content
Back to timeline
arXivSource publication:

AIMS turns natural-language deployment requests into sim-to-real ISAC configurations via a two-agent framework, improving vehicle detection and beam prediction on DeepSense 6G

Related research and updates

Synopsis

The authors propose AIMS, an agentic AI framework for sim-to-real multi-modal integrated sensing and communication (ISAC): given a natural-language deployment request specifying the target task, deployment conditions, and real-data budget, it derives a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model, using a two-agent architecture for scene construction and task learning; on the real-world DeepSense 6G dataset it improves vehicle detection and beam prediction over the considered simulation and fusion baselines, and a separate orchestration benchmark shows improved plan correctness with structured domain knowledge and validation feedback.

Source-provided article image: AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC
Fig. 1 ·

Fig. 1: Overview of AIMS for deployment-specific sim-to-real multi-modal ISAC learning. A natural-language deployment request is organized into a shared experiment state for two-agent planning and capability execution with validation-driven revision. The scene construction agent produces aligned sensing and wireless data, while the scene understanding agent configures learning and transfer to produce a deployment-specific task model.

arXiv

Interpretation

AIMS is an agentic framework that converts a natural-language deployment request (target task, deployment conditions, real-data budget) into a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model. Data-driven multi-modal ISAC models have depended heavily on annotated real-world data, and although synthetic data reduces that burden, adapting existing simulation pipelines to a target deployment requires consistent scene, sensing, wireless, and learning configurations whose mismatches impair sim-to-real transfer; AIMS automates this configuration and coordination. Described at the abstract level as a framework design, validated by experiments on the real-world DeepSense 6G dataset; configuration-derivation details and ablations are not expanded in the abstract.

A two-agent architecture is used: a scene construction agent generates geographically grounded, synchronized sensing and wireless records from shared physical states, while a scene understanding agent configures task-relevant modalities and mixture-of-experts (MoE) learning for zero-shot inference or few-shot adaptation. Decoupling scene construction from task learning into two cooperating agents lets sensing and wireless records share the same physical states and stay synchronized, while modalities and learning are selected per task. The abstract states the architecture and division of labor but gives no quantified per-component results.

Structured domain knowledge guides dependency-aware planning, and validation evidence supports feedback-driven revision of affected decisions. Domain knowledge serves as a planning prior and validation evidence as a replanning signal, so the agent makes dependency-consistent decisions across coupled configurations and can revise them. A separate orchestration benchmark evaluates task interpretation, dependency reasoning, and feedback-driven replanning, showing improved plan correctness with structured domain knowledge and validation feedback.

On the real-world DeepSense 6G dataset, AIMS improves vehicle detection and beam prediction over the considered simulation and fusion baselines. It grounds the sim-to-real multi-modal ISAC pipeline in two downstream tasks on a real dataset and reports improvements relative to baselines. Reported in the abstract as outperforming the considered simulation and fusion baselines on a real dataset; specific metric values, sample sizes, and statistical significance are not given in the abstract.

Perspective

The work targets sim-to-real transfer for multi-modal ISAC, applying to deployment requests that state the target task, deployment conditions, and real-data budget in natural language and producing a deployment-specific task model; the evaluation setting is vehicle detection and beam prediction on the real-world DeepSense 6G dataset plus a separate orchestration benchmark spanning diverse deployment requests. Its transferability points to similar deployment processes that must coordinate scene, sensing, wireless, and learning configurations.

The abstract does not give concrete metric values for vehicle detection and beam prediction, the real-data budget values, sample sizes, or statistical significance, nor does it separate the contributions of scene construction, scene understanding, MoE, domain knowledge, and validation feedback; the task-set size and scoring of the orchestration benchmark are also not expanded. These are open questions that require the main-text figures and experimental setup to resolve.

Sources