Engineering Sciences
140 items
WorldLine pretrains on 10,000 hours of action-free robot video and grounds 2,000 hours of action trajectories across ten embodiments, lifting robot-mask IoU by 0.1626 on failed trajectories and predicting trajectory success at 74% mean accuracy
WorldLine introduces an action-driven visual simulator that decouples transferable robot-object dynamics learning from heterogeneous action grounding: it pretrains on more than 10,000 hours of action-free robot videos, grounds them with over 2,000 hours of action trajectories across more than ten embodiments through image-space action maps, adds multi-view, failure-enriched and relational-regularized training, and distills a robot-focused few-step causal rollout; across held-out and out-of-domain settings it maintains visual quality and robot-motion agreement, improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, predicts trajectory success with 74% mean accuracy on RoboTwin and AgiBot, and improves task success by up to 21.
Real2Gym turns human demonstration videos into executable simulation gyms and trains a failed Franka task into success
Real2Gym is an agentic Real2Sim2Real framework that reconstructs human and robot demonstration videos into visually aligned, natively physics-validated Blender and MuJoCo interactive environments, where an agent generates executable code and distills successes and failures into reusable skills, reaching 87.5% task success with roughly 75% fewer policy-execution tokens than GPT-6 Astra across 24 reconstructed DROID and EgoDex environments and turning a zero-shot real-robot failure on narrow-clearance plate placement into success after simulation-based evolution on a Franka arm.
StructRL lifts long-horizon VLA success from 41.5% to 49.1% with verifiable subtask rewards
StructRL is an online reinforcement learning framework that automatically decomposes each long-horizon instruction into subtasks verifiable by binary environment-state criteria and organizes them into ordered dependency groups, granting intermediate rewards only after all prerequisites of a subtask are complete and scaling each reward by completion pace; across RoboCasa365 and LIBERO-Long with the GR00T-N1.5 and π0.5 VLA backbones it consistently outperforms the evaluated online RL baselines (e.g., on GR00T-N1.5, RoboCasa365 rises from 41.5% to 49.1% and LIBERO-Long from 92.4% to 96.6%).
EngiWorld tests 1,301 real engineering tasks and finds the strongest model reaches only 44.3 EngiScore, with 3.6% success on multi-software attempts
The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.
CrossBFM distills Unitree G1's latent behavior space onto three humanoids in under one GPU-hour, losing only 0.025 rad in tracking
CrossBFM freezes the backward map of a Behavior Foundation Model pretrained on the Unitree G1 and, using the frame-level cross-embodiment correspondence supplied by retargeted data, distills that latent space onto three humanoids (Inhouse M3, Booster T1, Fourier N1) with a unified encoder that has no robot-specific parameters, training the encoder in under one GPU-hour and each tracker in about 10 more GPU-hours; all three prompting modes transfer (motion tracking, goal reaching between poses, reward optimization), the latent-conditioned policy loses only 0.
Opus 5.5's agent-designed libraries cut downstream code and beat the human production library by 2.3 points, yet 11 of 15 tasks merely reproduce human abstractions
The authors introduce LibraryDesignBench, a two-phase benchmark in which a designer agent implements a full-featured library from a specification that lists required capabilities without prescribing interfaces, and three user agents from different model families then solve problems with it, scored by pass rate squared times simplicity relative to reference solutions written with the real production library; across 15 library-design tasks, 242 expert-validated problems and four languages, Opus 5.5 scores 48.9, 2.3 points above the production library, designers reproduce production-library abstractions on 11 of 15 tasks, 64% of audited excess code traces to rigid or hard-to-use interfaces rather than missing capabilities, and adding prescriptive agent-first guidance raises the score to 46.
EVO-WAM lets world action models self-train on their own generated video, lifting RoboTwin unseen-task success from 26.9% to 68.0%
The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
Tsinghua's Leap Lab finds in a matched comparison that world action models generalize from a single inference-time forward pass, not from denoising a clean future
Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
AnyStep-WAM distills frozen-teacher trajectories and schedules budgets by risk and benefit, cutting denoising steps by roughly half to 85% across three world-action models while holding success rates
The work introduces AnyStep World Action Model, a framework that performs budget-aligned flow-map distillation from frozen-teacher trajectory intervals and trains a lightweight risk-benefit scheduler to predict teacher-trajectory difficulty and budget-specific student fidelity from a single one-step preview, selecting the smallest denoising budget that meets a fidelity requirement; on RoboTwin 2.0 it reduces average denoising steps by 60.2%, 49.8%, and 85.28% on Motus, FastWAM, and LingBotVA while keeping average success within 0.24 percentage points of full-budget baselines, raises one-step success by 7.07, 12.08, and 8.94 percentage points respectively, and achieves 1.67-6.14x per-call speedups on six real-world manipulation tasks.
Page 3 · showing 10