Skip to main content
Back to timeline
arXivSource publication:

Architect-Ant trains on 505 real professional floor plans to produce editable furniture layouts, reaching 94% functional completeness with 0.6% collision and out-of-bounds rates

Synopsis

The work introduces AntPlan, a curated dataset of 505 real professional residential floor plans with dense annotations across 92 furniture classes and ten room categories, and Architect-Ant, which represents layouts in an editable coordinate DSL, adapts Gemma-4-31B-it via LoRA supervised fine-tuning, then applies GRPO with a room-aware Layout Rule Score (LRS) as final-layout reward; on 109 SceneSmith room shells it reaches 94% mean functional completeness, LRS 8.80, and 0.6% collision and out-of-bounds rates, outperforming the compared baselines while running at about 95.4 s per room, faster than iterative agentic approaches.

AI-generated editorial illustration: Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans

Interpretation

AntPlan provides 505 real professional architectural floor plans with structural, room, and furniture annotations spanning 92 furniture classes and ten residential room categories, averaging 53.01 annotated objects per plan. Compared with existing 2D floor-plan datasets such as CubiCasa5K's 10 coarse furniture classes or FloorPlanCAD's 30 classes, AntPlan annotates dense, fine-grained furniture directly on real professional drawings, substantially expanding furniture-class coverage. The paper tabulates MSD, ResPlan, ZInD, FloorPlanCAD, SESYD, CubiCasa5K, and AntPlan by scale, style, structural elements, room classes, and furniture classes; annotation combines an RT-DETR-X detector with manual review, and the authors explicitly treat reviewed furniture annotations as pseudo-labels.

Architect-Ant represents layouts in a coordinate DSL, first supervised-fine-tunes, then applies GRPO with LRS as final-layout reward, generating constraint-aware layouts without prescribed reasoning traces. Prior methods use layout data, explicit constraints, or pretrained model knowledge individually; Architect-Ant combines all three during training and turns geometric and functional constraints into outcome-level rewards rather than supervising intermediate reasoning steps. The paper reports full SFT and GRPO settings (LoRA rank 64, 5 epochs, 15 GRPO updates, 16 completions per room, 8 rooms per update for 128 rollouts) and compares base, SFT, and SFT+GRPO LRS on 243 held-out AntPlan rooms.

On 109 SceneSmith room shells, Architect-Ant achieves the highest mean functional completeness (94%) and highest LRS (8.80), with collision and out-of-bounds rates of 0.6% each, while producing the highest mean number of furniture objects per room. Against Holodeck, LayoutGPT, LayoutVLM, SceneSmith, and the learned ATISS, InstructScene, and DiffuScene baselines, it maintains low geometric violations without resorting to sparse layouts. The paper reports room-level means and specifies that COL+ applies category-aware filtering, OOB uses a surface-sampling test, and FUN+ counts room-specific functional slots; it also reports a 97.2% accessible-pair rate, within 2.3 points of the best method (Holodeck, 99.5%).

Two visual judges prefer Architect-Ant renders on the 109 shared room shells: Claude Sonnet 5 prefers it in 72 rooms and Kimi K3 in 74. This adds visual-preference evidence beyond geometric and functional metrics, and the aggregate LRS direction (8.80 vs. 8.44) agrees with the judges' overall preference. The paper describes anonymous A/B renders, randomized positions, and a shared prompt, and notes that these are automated visual preferences rather than a human user study, with one presentation order per pair leaving order sensitivity unmeasured.

Perspective

The result targets residential interior design, real-estate visualization, and 3D scene construction, and applies to four room types: bedroom, bathroom, kitchen, and living room. AntPlan and the trained adapters provide a starting point for extending to more object classes, richer geometry, and more detailed spatial relationships. LRS, as a rule-based score used for both training and candidate selection, can also serve as a reusable evaluation proxy for similar layout-generation methods.

LRS captures only the geometric and functional properties encoded by its rules; it does not measure aesthetics or guarantee full 3D usability. The authors report that in the kitchen reasoning branch the fraction of retained layouts with wall-overlap violations rises from 30.2% to 37.0%, and that the kitchen SFT+GRPO reasoning branch has a 54.0% excluded-sample rate, so high scores should be read alongside generation reliability. Reference layouts score a mean LRS of 8.73 with a median of 10, while Architect-Ant's 8.80 comes from a different room population, so it is not evidence of better-than-professional design quality. The visual-preference studies are automated, use one presentation order per pair, and the whole-house comparison draws on different floor plans and unequal selection pools. In addition, SceneSmith's 1,555 s entry is an object-count-normalized estimate rather than a measured rerun, so the timing comparison does not establish a hardware-normalized speedup.

Sources