Engineering Sciences
148 items
The Tasteful Agent: Measuring and Training Long-Horizon Judgment from Trajectory Hindsight
The work defines an agent's taste as its ability to choose the better direction before the outcome is visible, builds Taste-Bench, a 502-question benchmark mined automatically from decision forks in existing agent trajectories, finds that the best model answers only 59.7% correctly while forks whose deciding evidence appears later are much harder and a larger reasoning budget does not help, and shows that distilling the reasoning of a teacher that has seen the outcome into a student improves judgment on unseen tasks by 17.9 percentage points and raises end-to-end success on held-out SWE-bench Pro tasks from 14.6% to 33.7%.
A New Boron Allotrope, Imma-B60: Deformable and a Million Times More Conductive
Using a two-step route that first reacts boron with sodium under high pressure and then heats the mixture at 400 °C under vacuum to almost completely remove sodium, researchers prepared a new boron allotrope, Imma-B60, whose boron-atom network retains open space where sodium atoms once sat, enabling dislocation slip so the material stretches to 23% of its original length before breaking without springing back, and whose electrical conductivity exceeds that of typical boron materials by more than a million times.
TUM team surveys 329 works on physics-embedded robot learning: 71% encode physics in architectures, only 4% combine multiple embedding routes
This survey systematically reviews methods that embed physics priors into robot learning and proposes a unified taxonomy classifying existing work by where physics is embedded—physics-guided inputs, data, and representations; physics-encoded model architectures; and physics-informed training losses—then reviews methods for robot dynamics learning, trajectory planning and prediction, control, and estimation together with the open-source software ecosystem, reporting that physics-encoded architectures account for 71% of surveyed methods, physics-guided 16%, and physics-informed 13%, while only 4% combine more than one embedding route.
RoboFollow diagnoses nine embodied policies in high-entropy scenes: near-saturated L0, sharp L1–L3 intent drops, and no fix from stronger VLMs or existing optimizations
The work introduces RoboFollow, a diagnostic benchmark combining high scene entropy, an L0–L3 hierarchical perturbation protocol, and stage-wise Intent/Execution scoring, and finds across nine VLA and WAM policies that strong in-distribution L0 performance does not transfer to L1–L3, with stronger VLM backbones, QA co-training, LangForce, and classifier-free guidance all failing to close the gap.
Agensh: Scaling Organizational Intelligence to 1,024 Agents
Agensh introduces a self-organized multi-agent harness without a central orchestrator, in which concurrent workers run an asynchronous cooperation loop supported by a shared workspace, a message interface, and shared context; on the five hardest ProgramBench tasks it raises the mean final test-pass rate from 19.31% with 1 agent to 28.78% with 128 agents, and on pandoc from 33.89% with 1 agent to 55.06% with 1,024 agents.
NVIDIA used an AI coding agent to migrate a Depth Anything 3 ROS 2 node to the CUDA buffer backend, letting depth images move between nodes with zero-copy transport
This NVIDIA tutorial shows how an AI coding agent using the migrate-node-to-rosidl-buffer skill migrates an already GPU-accelerated Depth Anything 3 TensorRT ROS 2 node to ROS 2 Lyrical's rosidl::Buffer and NVIDIA's CUDA buffer backend, so the output depth image's Image.data field is backed by CUDA storage and can move between nodes without serialization or host copies when conditions such as the same host, CUDA device, Linux user, and a supported RMW implementation are met, while preserving standard message interfaces and a CPU fallback path.
HuRo turns five human-video sources into about 142 million robotized frames, lifting real-robot OOD completion from 34.9% to 72.2% in ALLEX VLA pretraining
HuRo introduces a robotization pipeline that converts heterogeneous egocentric human videos into robot-aligned observations and action trajectories, builds the HuRo dataset of about 142 million frames from five sources (EgoDex, EgoVerse, Ego4D, Ego10K, and EPIC-Kitchens), and uses it to pretrain VLA policies on the ALLEX bimanual dexterous robot: as robotized data increases, average completion across four real-world manipulation tasks rises from 51.5% without pretraining to 80.3% with the full dataset, with ID completion from 68.1% to 88.4% and OOD completion from 34.9% to 72.2%.
CARE turns robot execution failures into training data, lifting average task success by 14.5 points in simulation and 15.9 points in the real world
The work proposes CARE, which collects failed rollouts, models stage-conditioned post-failure geometric deviation distributions, synthesizes representative failure states and corrective demonstrations from them, and at inference combines stage-wise planning with physically grounded 3D point-cloud monitoring to trigger atomic adjustments or re-operations; across RoboTwin 2.0, RoboFactory, the newly introduced FSR-Bench, and real-world dual-arm tasks it reports average task-success gains of 14.5 points in simulation, 15.9 points in the real world, and a 7.5-point gain in recovery success on FSR-Bench.
TAPe+ML v3 reaches 84.7 mAP50 and 65.3 mAP50-95 on COCO detection with under 100K parameters, and lifts 30-image industrial ore-blockage accuracy from about 28% with YOLO 26s to about 65%
The work presents TAPe+ML v3, which replaces raw pixel tensors with a structured TAPe (Theory of Active Perception) representation and combines background, pointer, and prototype-clustering submodels under a coordinator; with fewer than 100,000 parameters it reports 84.7 mAP50 and 65.3 mAP50-95 on COCO detection, 80.7 mask mAP50 and 58.4 mask mAP50-95 on COCO instance segmentation, 88.1% ImageNet-1k Top-1 and 89.9% ImageNet-Real, 92% versus 47% on Imagenette under identical training with TAPe versus raw pixels, and an industrial ore-blockage pilot where backbone adaptation outperforms head-only adaptation.
Page 11 · showing 10