Skip to main content

AI Core

989 items

  1. arXiv

    HiRAE fuses all 24 DINOv3-L layers with depth-grouped residual budgets, cutting ImageNet-256 reconstruction FID from 0.299 to 0.209 and lifting post-fine-tuning GenEval to 87.70

    HiRAE introduces a hierarchical representation autoencoding framework that groups all 24 layers of a frozen DINOv3-L encoder by depth into shallow, middle, and deep groups, learns residual corrections to the deepest representation under group-wise norm caps with tighter budgets for shallower groups, and jointly trains the fusion module and decoder while preserving the latent token count and channel dimension, reducing ImageNet-256 reconstruction FID from RAEv2's 0.299 to 0.209, raising PSNR from 22.667 to 26.377 dB and lowering LPIPS from 0.074 to 0.043, lowering guided generation gFID from 1.060 to 1.038, and improving text-to-image alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning, with post-fine-tuning GenEval rising from 84.
  2. arXiv

    ZJU team runs 25 same-family on-policy distillation pairs on Qwen2.5 from 0.5B to 14B, finding peak capability is predictable from student scale and teacher score, and that weaker teachers can teach stronger students

    The study systematically characterizes the scaling properties of same-family on-policy distillation (OPD) on Qwen2.5 (0.5B–14B) math reasoning, finding a regular early useful-transfer regime in which held-out accuracy rises approximately linearly with the square root of token-level reverse KL from the student initialization, and fitting power laws in student scale, teacher scale, and teacher gold score that predict peak accuracy and transfer rate, with every weak-to-strong student peaking above its own teacher.
  3. arXiv

    PreviewDiff puts a multimodal critic inside the denoising loop, beating Best-of-N at matched compute on SDXL and LTX-Video

    PreviewDiff introduces a training-free test-time search method that decodes partial previews at selected denoising checkpoints, asks a multimodal judge to score and critique them, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations that are scored and selectively rolled forward, consistently outperforming budget-matched Best-of-N and scalar-search baselines on image and video generation benchmarks.
  4. arXiv

    HSTA tracks technology diffusion across 30,000 arXiv preprints and USPTO patents, finding the highest semantic drift in Large Language Models (0.332) while paper-volume velocity fails to Granger-cause frontier compute surges

    The study introduces Hyperspherical Semantic Trajectory Analysis (HSTA), an unsupervised pipeline that encodes 20,000 arXiv preprints and 10,000 USPTO patent abstracts with Sentence-BERT, projects them onto a unit hypersphere, clusters them into eight sub-topics with Spherical K-Means alongside UMAP reduction, and defines two metrics, Semantic Centroid Vector Drift and Commercialization Offset, then links them to Epoch AI compute data through Vector Autoregressive Granger tests, finding that Large Language Models (drift 0.332) and Artificial Intelligence Systems (0.234) evolve fastest semantically while quarterly paper volume velocity alone does not Granger-cause frontier training compute surges at conventional significance levels.
  5. arXiv

    PrismQuant rotates dominant activation energy into the null space of grouped quantization, reaching 3.85 perplexity at W4A4KV4 on Llama-3.1-70B, only 0.22 points below full precision

    The work introduces PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the group-constant subspace of asymmetric grouped INT4, formulates rotation design as a Ky Fan trace maximization with a provably optimal closed-form solution, realizes it with compact Householder (compact-WY) transforms without gradient training, and evaluates W4A4KV4 post-training quantization on Llama, Qwen, and Mistral (including dense models up to 70B and a 30B mixture-of-experts model): on Llama-3.2-3B it sets the state of the art among compared methods in perplexity and accuracy, on Llama-3.1-70B it attains 3.85 perplexity and 72.46% average zero-shot accuracy (0.22 percentage points below full precision), and in a Llama-3.
  6. arXiv

    WISE-ATTA shifts active test-time adaptation from which sample to label to when to label, reaching lower error on ImageNet-C/R/K/A with fewer labels

    The work introduces budgeted active test-time adaptation (ATTA), in which labels are available for only a fraction of batches in a long test stream, and proposes WISE-ATTA: lightweight online signals (the fraction of low-entropy predictions) plus a sliding-history quantile threshold and a budget-debt controller decide when to request supervision, while prediction drift relative to an EMA anchor selects a single sample to label; on ImageNet-C under CTTA/FTTA and on ImageNet-R/K/A it matches or improves on recent ATTA methods at one label per batch and reports up to 50% fewer labels.
  7. arXiv

    EVO-WAM lets world action models self-train on their own generated video, lifting RoboTwin unseen-task success from 26.9% to 68.0%

    The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
  8. arXiv

    Omni-Decision turns the planning bottleneck of omni-modal agents into an attributable object, reaching 81.4% on OmniGAIA at about 43% of Gemini-3.1-Pro's cost

    The work presents Omni-Decision, which replaces the growing dialogue history with a task-scoped evidence ledger: a critic digests each noisy multimodal observation into typed evidence events committed by a deterministic reducer, so the planner always decides over a compact, verified state; controlled backend replacements show that replacing the planner costs far more than replacing the perception backend, and the system reaches 81.4% accuracy on OmniGAIA at roughly 43% of Gemini-3.1-Pro's per-question cost and 65.0% on WorldSense long-video understanding, with execution trajectories used to fine-tune and reinforce weaker planners.
  9. arXiv

    VideoLoop uses dual-loop bounded working memory to ease semantic thrashing in long-video agents, lifting Gemini 3.1 Pro by 3.2 to 4.5 points on VideoMME (long) and two other benchmarks

    The work formalizes a failure mode it calls semantic thrashing, in which append-only working memory keeps accumulating noise and dilutes key evidence, and proposes VideoLoop: an outer multimodal agent explores the video inside a sandboxed filesystem while an inner memory orchestrator retrieves relevant artifacts from an unbounded filesystem and rewrites a bounded working memory after every step; across VideoMME (long), VideoMMMU, and LongVideoBench (long), VideoLoop improves four LVLM backbones in a plug-and-play manner, averaging a 4.2-point gain over baseline on VideoMME (long) and reaching 88.3%, 88.8%, and 80.9% with Gemini 3.1 Pro.
  10. arXiv

    One token-level interaction decomposition links data selection, collision-versus-erosion forgetting, and plasticity loss into a single learning-dynamics account

    The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.

Page 27 · showing 10