Engineering Sciences
148 items
Splitting repository-level SWE tasks into category experts and distilling them back into one model lifts Pro-618 mean resolution to 58.04%
Addressing a "category see-saw" in repository-level software engineering, where pooled agentic reinforcement learning improves some task categories while regressing others and aggregate resolution hides the change, the work builds a category-aware expert-training and policy-integration framework: SWE Labeler groups tasks by repository domain into service/data/security (A), user-facing applications (B), and systems/tooling/runtimes (C); Agentic-miniRL with a Refresh–Repair–Expand loop trains same-origin category experts; and label-routed multi-teacher on-policy distillation (MOPD) consolidates them into one deployable student, which reaches 58.04% mean resolution on Pro-618 (+5.39 points over base) and 59.00% on SWE-bench Multilingual (+2.
Aligning to Frozen World-Model Features Lifts a 0.8B Robot Policy to 97.9% on LIBERO and +2.3 Points on RoboCasa-GR1 at Zero Inference Cost
The work introduces a method that distills world-model representations into compact vision-language-action (VLA) policies: ordinary VLA training gains one cosine alignment term against cached features from a frozen world model, the teacher runs only once offline and is never loaded during training, the projector is discarded afterward, and the deployed policy is identical to the undistilled baseline; the 0.8B student reaches 97.9% average success across the four LIBERO suites (2.6 points above the identical undistilled student), improves RoboCasa-GR1 humanoid manipulation from 48.2% to 50.5%, and gains on both single-arm and bimanual real platforms, with the gain surviving changes of student scale, backbone, alignment layer, and teacher.
Reading Without Unrolling: Virtual Unwrapping and Lead-Ink Imaging of Herculaneum Papyri
Together, an external news story and two papers present two independent routes: one tests whether imaging lead rather than carbon in the ink can make text legible, using home-made burnt papyri, while the other uses synchrotron phase-contrast micro-CT plus machine learning to achieve the complete virtual unwrapping and reading of an intact, unopened Herculaneum roll, PHerc. 1667, recovers title and author evidence identifying PHerc. 139 as Philodemus, On Gods, Book 8, and provides volumetric ink validation in PHerc. Paris 4.
Scoring the scorer: mutation analysis finds the official GPU-kernel check misses one in six witnessed faults, with 78.6% of precision faults escaping
The work adapts mutation analysis into an adequacy metric for graded numerical oracles, injecting 10,303 compilable faults into verified CUDA implementations of 188 KernelBench problems (7,384 with independent kill witnesses), measuring that the official check deterministically misses 16.9% of witnessed faults with misses skewed by family (8.7% arithmetic, 78.6% precision), and using that measurement to explain mechanisms, audit existing patches, and synthesize two-input suites reaching 98.0% detection.
ShieldVLA gates VLA fine-tuning with Hamilton-Jacobi reachability, cutting cumulative safety cost by 57% and lifting success rate by 0.13
ShieldVLA is a safety-aligned fine-tuning framework for vision-language-action (VLA) models that learns a model-free Hamilton-Jacobi reachability value function from visual observations as a safety critic, uses that critic to gate policy optimization into reward maximization inside feasible regions and recovery near unsafe states, and replaces manual per-step cost labels with rubric-based VLM safety scores; across five navigation and manipulation benchmarks (Dubins-VL, TurtleBot-Nav, Safety-CHORES Nav/Fetch, Franka-Reach) and multiple VLA backbones it reduces cumulative safety cost by 57% on average and improves task success rate by 0.13 over SafeVLA.
GAM swaps language and video pretraining for a frozen 3D grounding backbone, lifting RoboTwin 2.0 randomized-scene success from 30.4% to 47.6%
The work proposes Grounded Action Model (GAM), which builds a robot foundation model on a pretrained promptable 3D grounding model (WildDet3D), maps language, 2D point, or 2D box prompts into a shared object-centric representation (image tokens keeping target visual features plus detection tokens encoding point clouds, metric positions, and extents), and uses a multi-stream transformer (MM-DiT) with robot state history to predict action chunks; it reaches 55.3% average success across 50 RoboTwin 2.0 tasks (vs. 52.0% for Spatial Forcing), 47.6% under scene randomization (vs. 30.4% for Abot-M0), a state-of-the-art 61% average across 16 LIBERO-PRO perturbation settings (vs. 53%), 17/20 successes under visual shift on a bimanual YAM (vs. 4/20), and 64.7% in-distribution and 49.
Ultra-cold microscopy reaches the lab: the liquid-helium STEM GAIA and the road to atomic-resolution cryogenic imaging
This evidence bundle centres on a Nature news feature reporting that Bruker will ship GAIA, described as the first scanning transmission electron microscope to operate stably near absolute zero with liquid helium, to three laboratories in Canada, Germany and the United States, where it can image materials at about −266 °C (7 kelvin) for 30 hours or more, while tracing the earlier route: a 2025 US-reported TEM attachment reaching −253 °C (20 kelvin) with atomic-resolution imaging stable for more than ten hours, a condenZero liquid-helium attachment at around 4 kelvin for 24 hours that did not aim for atomic resolution, and Idrobo's 2021 comment on aberration correction combined with further instrument developments.
DeformSmith: Turning Text or a Single Image into Graspable Deformable Assets via a Physics Harness
DeformSmith presents a hierarchical agentic framework that, from text or a single image, progressively builds geometry, a physical model, material behavior, and robot interaction, using a shared physics-grounded harness to evaluate and revise candidates through simulation probes and robot pick-and-place feedback; across 39 cases its assets score higher on visual quality and physical plausibility than baselines including PhysGen3D, PhysGM, and PhysX-Omni, while also producing replayable manipulation data.
Foresight Without Trajectories: Guiding 3D Diffusion Policies with a Latent Movement Trend
The work introduces Movement Trend Guidance: a compact latent of interaction evolution is learned from a short history of point clouds and robot states, supervised during training only by sparse future gripper states, and at inference retained alone as future-oriented conditioning, consistently improving 3D diffusion policies on RoboTwin2.0, LIBERO-40, DexArt and five real-robot tasks while adding only 3.52% more parameters than DP3.
Page 12 · showing 10