AI Core
977 items
EngiWorld tests 1,301 real engineering tasks and finds the strongest model reaches only 44.3 EngiScore, with 3.6% success on multi-software attempts
The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.
ANTMAN replaces static partitioning with a revisable Need Graph: a 16x larger search space raises active coordination only 1.23x versus over 15x for partition-driven baselines
The work introduces ANTMAN, a multi-agent information-seeking framework that treats evolving unresolved information needs, rather than input partitions, as the unit of runtime coordination, maintaining a revisable Need Graph that tracks unresolved requirements, accumulated evidence, prior attempts, and search progress to control worker selection, routing, and task-local recovery; across multi-document QA, controlled long-context scaling, and structured navigation, ANTMAN increases active coordination by only 1.23x under a 16x increase in searchable context, compared with more than 15x for partition-driven baselines, while preserving strong answer quality and retaining most of its performance when execution is delegated to substantially smaller 8B worker models.
Porcedda's Sys1Cal-v1 shows Jev's Choice probabilities are systematically distorted, and restoring an uncertainty component lifts median soft accuracy from 0.771 to 0.978
Porcedda built Sys1Cal-v1, a synthetic dataset of 365 True/False items derived from 92 probability problems whose proposition probabilities are known by construction, and used total variation distance and distributional overlap (soft accuracy) to evaluate Jev's Noul, Choice, and Score primitives plus the open-source baseline SemIf, finding that Jev's Choice probabilities are systematically distorted and that inverting a latent third truth value of uncertainty estimated from Score expectations raises median Choice soft accuracy from 0.771 to 0.978.
LEGO-Anything has coding agents rebuild 3D scenes from a single image as Blender code, with GPT-6-astra scoring 53.4% indoors and 39.6% outdoors and LEGO-Plugin lifting all six models by up to 62.7%
The work presents LEGO-Anything, an Image-to-Code framework in which a general-purpose coding agent reconstructs a 3D scene from a single image by iteratively writing, executing, and revising Blender programs, yielding an executable, inspectable, editable, and queryable scene program, together with LEGO-Bench, a simulator-grounded benchmark of 208 images rendered from 104 indoor and outdoor scenes that separately scores artifact validity, visible-surface geometry, and rendered appearance; among evaluated agents GPT-6-astra achieves the strongest overall results (53.4% indoor, 39.
ALICE estimates mutual information zero-shot with one pretrained Transformer, cutting error at least twofold at the 1k-sample budget on the Beyond Normal benchmark and reproducing prior findings in biology, genetics, and neuroscience data
The authors present ALICE, a foundation model trained once on a synthetic distribution family that turns mutual information estimation into in-context estimation of rectified-flow velocity fields for unseen distributions, matching per-distribution neural estimators on the Beyond Normal benchmark with at least twofold lower error at a 1k-sample budget and zero-shot reproducing and extending dedicated estimators' findings on biology, genetics, and neuroscience data the model never saw.
Quantizing softmax inside attention: whether pretraining works from scratch depends on the backward rule, with detached row-extrema gradients diverging late and MinMax plus Weight-STE trailing softmax by 0.89 nats at 2.5B tokens
This work studies replacing exact softmax with a quantized softmax during pretraining (K-interval attention, which approximates the exponential with K+1 grid values), derives the corresponding backward rules including calibration derivatives, and compares per-row grid calibration (MinMax versus a fixed window, FWM), interpolation (LERP) versus hard rounding (Nearest), and placement of a straight-through surrogate before normalization (Weight-STE) or after it (Prob-STE) in pretraining experiments matched on model, data and optimizer; it finds that detaching the row extrema leaves the forward unchanged but causes a delayed divergence after 25–30M tokens ending 0.65–3.07 nats above the matched run, that under hard rounding MinMax with Weight-STE ends 0.
Raven composes model–harness pairs into a multi-agent ecosystem and leads planning, four specialist domains, and skill reuse over Claude Code and other baselines
Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
LIFT controls video generation with a last-frame layout plus camera trajectory, raising mIoU from 0.41 to 0.51
LIFT introduces a unified image-to-video framework that adds Layout-In-FuTure control on top of camera control, letting users specify what should appear in a future view and where; it transfers a dense per-frame layout teacher to a last-frame-layout student via dual-mode on-policy self-distillation (OPSD) and curates the LIFT-Vista dataset for large viewpoint changes, reporting improvements in video quality, future-layout controllability, and camera controllability over compared methods.
FRAC swaps exponential forgetting for power-law long memory in SSMs, beating Mamba and GDN on long-context 1.3B language modeling
The work introduces FRAC, a selective state space model architecture derived from fractional dynamics that approximates a heavy-tailed fractional kernel with a finite-state, log-spaced sum of exponential modes, replacing exponential forgetting with power-law long memory and improving long-context performance over SSM baselines such as Mamba2, GDN, and Mamba3 on synthetic long-tail and recall tasks, 1.3B-parameter language modeling, and DNA modeling, while staying competitive on short-context tasks.
Page 24 · showing 10