Skip to main content

AI Core

998 items

  1. bioRxiv

    SpatialTRACE extends sparse annotations into tissue-wide anatomical axis and region maps and predicts the same coordinates from DAPI images alone

    The authors developed SpatialTRACE, comprising graph- and image-based models: SpatialTRACE-Graph combines gene-expression profiles with spatial neighborhoods to predict crypt-villus and epithelial-distance axis coordinates and to identify Peyer's patches in mouse small-intestine sections using as few as 10 annotated training villi, while SpatialTRACE-Image, a multiscale vision transformer that learns from the coordinate and region predictions generated by SpatialTRACE-Graph, predicts the same anatomical axis coordinates and regions across entire tissue images from DAPI alone and was applied to immunofluorescence images to map antigen-specific P14 CD8 T cells responding to acute systemic LCMV Armstrong infection in the small intestine, showing that a retinoic acid receptor inhibitor-treated
  2. Journal of Machine Learning

    OptimAI turns natural-language optimization problems into solver code with a multi-agent LLM pipeline, reaching 88.1% on NLP4LP and 82.3% on Optibench and cutting error rates by 58% and 52% over the prior best

    The work introduces OptimAI, an LLM-powered multi-agent framework that takes a natural-language optimization problem through four stages—formulation, planning, solver code generation, and reflective debugging—and adds UCB-based debug scheduling to switch dynamically among candidate plans; under zero-shot prompting it reaches 88.1% accuracy on NLP4LP with GPT-4o+o1-mini and 82.3% on Optibench with DeepSeek-R1, reducing error rates by 58% and 52% over the prior best, while ablations show that removing the planner or code critic drops productivity by 5.8× and 3.1× and that enabling UCB debug scheduling adds a further 3.3× productivity gain.
  3. bioRxiv

    DeepFisFis processes 5 ms audio segments in about 2.5 ms, detecting mouse ultrasonic vocalizations in real time and triggering closed-loop stimulation

    The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
  4. Journal of the Physical Society of Japan

    Hasegawa and Ohzeki compare Ising and QUBO encodings under fixed model, sampler, and step size, finding QUBO has lower Fisher spectral entropy and slower SGD convergence while natural gradient descent makes the encodings converge alike

    Under a controlled protocol that fixes the Boltzmann machine model, simulated-annealing sampler, and learning-rate design, the study compares Ising ({−1,+1}) and QUBO ({0,1}) variable encodings, exploits the identity that the Fisher information matrix equals the covariance of sufficient statistics to visualize empirical moments, and finds that QUBO induces larger cross terms between first- and second-order statistics, creating more small-eigenvalue directions and lowering spectral entropy, which explains slower convergence under stochastic gradient descent, whereas natural gradient descent, which rescales updates by the Fisher information matrix metric, achieves similar convergence across encodings due to reparameterization invariance.
  5. bioRxiv

    Murmurent layers agentic AI beneath lab collaboration and is used to seek putative Pin1 inhibitors

    The authors present and open-source Murmurent, shared software that sits beneath agentic AI for biomedical labs, offering multi-member project and "choreography" infrastructure, specialized agents for typical biomedical data science tasks, tiered memory, traceability records, SOP and data-governance enforcement, and multi-user collaboration, and they use the system to identify putative inhibitors of Peptidyl-prolyl cis-trans Isomerase NIMA-interacting 1 (Pin1), describing several approaches and the results they yield.
  6. arXiv

    G²PTQ refreshes gradients and Hessians before quantizing each Transformer block, halving KL divergence versus GPTAQ and adding 6.45% QA accuracy at 2-bit

    The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
  7. arXiv

    Frozen-base logit correction: CRN v2 fixes 53.3% of errors with no MMLU/BoolQ degradation, while LoRA fixes 83.3% but loses 17–75 points

    The study introduces CRN v2, a 34M-trainable-parameter (0.73% of the 4.65B text module) logit-level correction module sitting atop a fully frozen Gemma 4 E2B, trained by supervised fine-tuning plus reference-free DPO on 83,400 error-correction pairs, which corrects 53.3% of base-model errors on a 60-question CEHRI domain exam (43.3% on a reworded variant) with no degradation on MMLU, BoolQ, or car-wash, whereas a parameter-budget-matched LoRA baseline corrects 83.3% but loses 30–75% capability on the same benchmarks.
  8. arXiv

    CompoWorld composes 448 reusable services into cross-service tasks, lifting Qwen3.6-35B-A3B by 9.17 points on average across eight benchmarks

    CompoWorld introduces compositional environment scaling: coding agents turn MCP tool specifications into executable services with typed states and shared interfaces, a random walk connects services through dependency graphs to generate and verify cross-service tasks, verified trajectories support SFT while a Completion-Focused Rubric Reward guides GRPO-based RL, and with 448 services exposing 10,130 tools plus 3K SFT trajectories and 1K RL tasks, Qwen3.6-35B-A3B improves by 9.17 points on average across eight benchmarks and raises AutomationBench task success from 10.33% to 32.33%.
  9. arXiv

    Rolling-WAM spreads joint denoising across replanning cycles, reaching 98.1% on LIBERO and 93.3% on RoboTwin 2.0 with a 4.5x steady-state replanning speedup over standard joint WAMs

    Rolling-WAM introduces a rolling world action model that keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk each replanning cycle while partially refining farther-future chunks, thereby distributing joint denoising computation across control cycles; it achieves competitive manipulation success on LIBERO, RoboTwin 2.0, and real-world Unitree G1 humanoid tasks while delivering a 4.5x steady-state replanning speedup over standard joint WAMs.
  10. arXiv

    SolveEdit's 2,728 scene-transformation tasks leave the best model at 57.0% SolveScore, while a two-stage planner lifts GPT-Image-2 to 71.6%

    The authors introduce SolveEdit, a benchmark that formulates visual problem solving as inferring and executing a valid scene transformation from a given image and goal while preserving unrelated content, comprising 2,728 cases across 10 domains and 54 subdomains organized by whether the required transition is fixed by the instruction (IS), the scene state (SD), or an in-image rule (RD), and scored by atomic transition contracts and SolveScore without a single reference output; across nine image-to-image and two image-to-video models the strongest reaches only 57.0% SolveScore, with RD trailing IS by 17.2 to 23.

Page 32 · showing 10