Skip to main content

Daily report

AI and science frontiers · 2026-09-25

Only content delivered through the publication boundary on this date is included.

Amazon Science

A kernel-centric path to real-time video generation on Trainium

Using Rolling Forcing as a proxy for autoregressive diffusion video generation, Reactor and the Amazon Neuron Science team adopted a kernel-centric, bottom-up approach—writing hardware-level kernels via the Neuron Kernel Interface, applying hybrid sequence and tensor parallelism sharding, and rewriting model code to batch shared components of the diffusion and cache-update phases—successfully generating a correct video on the first end-to-end run on Trainium and delivering real-time generation above 16 fps.
Scientific Reports

LLM pairwise comparison of historical proposals at three Spallation Neutron Source beamlines ranks them in positive correlation with human ranking, matches human reviewers at flagging high-publication-potential proposals, and costs over two orders of magnitude less

Using historical general user proposals from three beamlines at Oak Ridge National Laboratory's Spallation Neutron Source (EQ-SANS, CNCS, POWGEN), the study has large language models judge every pair of proposals within a run cycle, converts the win-lose outcomes into rankings with the Bradley-Terry model, and finds that LLM rankings correlate positively with human rankings (Spearman ρ about 0.2-0.8, rising to ≥0.5 after 10% outlier removal), show no statistically significant difference from human review in identifying proposals with high publication potential, cost over two orders of magnitude less, and support linear-complexity proposal similarity analysis via embedding models.
Nature Communications

StriMap integrates physicochemical, sequence-context, and interface-structural features to predict TCR–peptide–HLA recognition, and screening 13 million peptides from 43,241 bacterial proteins yielded candidate molecular mimics that activated T cells expressing an ankylosing spondylitis-associated TCR

The work presents StriMap, a unified framework that predicts TCR–peptide–HLA interactions by integrating physicochemical, sequence-context, and structural features at recognition interfaces, reporting state-of-the-art performance with improved generalizability; as a case study, the authors screened 13 million peptides from 43,241 bacterial proteins and identified candidate molecular mimics that were experimentally validated to activate T cells expressing an ankylosing spondylitis (AS)-associated TCR, with a top validated peptide enriched in patients with inflammatory bowel disease (IBD), suggesting potential shared microbial triggers.
发表出处待核验

400 ratings from 20 music-domain experts show LLM judges align only moderately with humans (up to r=0.55) yet far exceed BLEU and ROUGE-L reference baselines

This work presents the first user study assessing the reliability of LLM-as-a-judge for evaluating conversational recommendation system (CRS) responses: 20 multi-turn sessions were sampled from TalkPlayData-Challenge, candidate responses were generated by four instruction-tuned LLMs, and 20 music-domain experts produced n=400 ratings on Personalization Quality and Explanation Quality; bootstrapped correlation analysis over 10,000 iterations found moderate positive alignment between LLM judges and human assessments (highest r=0.55 for Gemini-3.1Pro on personalization), outperforming all reference-based baselines.
International Journal of Robust and Nonlinear Control

Composite adaptive control barrier functions couple parameter estimation with safety in one energy function, recovering in three simulations the safe operating space robust methods give up

This paper presents the composite adaptive control barrier function (CaCBF) algorithm for nonlinear control-affine systems with linear parametric uncertainty, deriving the adaptation law from a composite energy function that integrates a logarithmic safety barrier, a control Lyapunov function, and a parameter-error term; it proves forward invariance of the safe set for all bounded parameters without persistence of excitation, robustness of the safety guarantee to bounded state-derivative measurement errors, uniform ultimate boundedness of all closed-loop signals, and that the CaCBF admissible control set always contains the robust counterpart as a subset; simulations of adaptive cruise control, an omnidirectional robot, and a planar drone traversing a narrow gate show CaCBF recovers the pe
arXiv

IntraStyler uses a contrastively trained style encoder to synthesize T2 MRI styles without predefined sub-domains, raising Extra-VS Dice from 0.61 to 0.81 with zero failures on CrossMoDA

The work proposes IntraStyler, a 3D unpaired image translation method that requires no predefined sub-domains, using a contrastively trained style encoder to extract anatomy-disentangled style embeddings from each target-domain image and conditioning a QS-Attn synthesis network via dynamic instance normalization; on the CrossMoDA benchmark (226 labeled ceT1, 295 unlabeled T2, 96-image T2 test set) it automatically discovers 7 style references instead of the 3 used by sub-domain methods, and downstream nnU-Net segmentation achieves the best Dice and ASSD on Intra-VS, Extra-VS, and Cochlea, with Extra-VS Dice of 0.81 and zero failures.
arXiv

Treating electrode coordinates as learnable parameters: percept-aware optimization on folded cortex improves reconstruction fidelity while eliminating vascular safety-margin violations

The study presents a percept-aware surgical planning framework for cortical visual prostheses that treats 3D electrode coordinates as learnable parameters and optimizes them end-to-end through a differentiable forward model of prosthetic vision on FreeSurfer fsaverage folded cortical geometry, minimizing task-level perceptual error subject to vascular avoidance and gray matter feasibility constraints; on simulated MNIST reading and CIFAR-10 natural image tasks it consistently improved reconstruction fidelity over visual field tiling and visual field coverage baselines (median MSE reductions of up to 67.7% and 33.4% on MNIST, with downstream classification accuracy gains of 62.6% and 22.4%), eliminated all 300 µm safety-margin violations with only 1.7% (MNIST) and 4.
arXiv

SEA-PEFT raises mean Dice by 2.4–2.8 points in 1/5/10-shot 3D medical segmentation while training under 1% of parameters

The authors propose SEA-PEFT, which treats adapter configuration as an online allocation problem solved during fine-tuning through a search–audit–allocate loop that trains active adapters, estimates each adapter's Dice utility by momentarily toggling it off, and reselects the active set under a parameter budget with a greedy knapsack allocator, stabilized by EMA+IQR smoothing and a finite-state machine; on TotalSegmentator and FLARE'22 it improves mean Dice by 2.4–2.8 points over the strongest fixed-topology PEFT baselines across 1/5/10-shot settings while training under 1% of parameters.
发表出处待核验

Google Deploys a Dual-Loop LLM Agent System on YouTube That Autonomously Found RMSprop, Gated Path, and a Multi-Objective Reward, Lifting Multiple A/B Metrics

The work proposes and deploys a self-evolving recommendation system in which Gemini-family LLM agents act as machine learning engineers: an Offline Agent (Fast Loop, waking every 5 minutes and running daily) generates and offline-scores candidate configurations, while an Online Agent (Slow Loop, running daily) ranks candidates and drives A/B experiments; across several YouTube recommendation surfaces the agents autonomously discovered optimizer switches (Adagrad to RMSprop, FTRL), a Gated Path (GLU) architecture, activation refinement, multi-objective reward synthesis, and reward hyperparameter tuning, with multiple online metrics improving at the 95% confidence level and experiment throughput rising from Θ(1)-Θ(10) per week to Θ(100) per week.
arXiv

TSegAgent achieves zero-shot tooth segmentation and FDI identification by pairing SAM3 with geometric reasoning, reaching mIoU 93.37 and TIR=1 of 87.17% on Teeth3DS while generalizing to a private dataset

The work proposes TSegAgent, which reformulates tooth instance segmentation and FDI labeling of intra-oral scanned 3D models as a zero-shot geometric reasoning problem: multi-view renderings with curvature heatmaps and the SAM3 text prompt "tooth" produce candidate masks, which are merged into face-level instance labels via IoU and containment relations, after which a geometry-aware vision-language agent performs non-tooth region identification, central incisor localization, full-arch classification, and error correction through multi-round conversation; it reports mIoU 93.37, TLA 96.40, TSA 96.76, TIR 97.46, and TIR=1 87.17 on Teeth3DS with 1200 3D tooth models, and mIoU 82.10, TLA 96.68, TSA 95.99, TIR 85.41, and TIR=1 51.
Anthropic

Von Hippel's nine-loop bounty: Anthropic's Claude computed the six-particle nine-loop amplitude for about $100 of compute, and Dixon independently validated it

Physicist and science writer Matt von Hippel publicly challenged AI companies to solve a hard scattering-amplitude problem on an academic budget; Anthropic's Liam Fitzpatrick and Siddharth Mishra-Sharma used Fable 5.1 inside the Claude Science platform, gave Claude a single problem statement and then mostly just told it to keep going, and Claude computed the six-particle (hexagon) nine-loop amplitude in planar N=4 super Yang-Mills two different ways, via the original bootstrap and via the indirect form-factor route, with the bootstrap portion costing about $100, equivalent to running 96 CPUs for a week; Lance Dixon of SLAC then independently validated the result, and Song He's group at the Chinese Academy of Sciences had also computed the symbol piece of the nine-loop amplitude with GPT-6
SIAM Journal on Scientific Computing

FLAME-derived LTLT factorization algorithms for skew-symmetric matrices, with fused BLAS-like operations, greatly outperform PFAPACK and Pfaffine

This work systematically derives a family of algorithms for the LTLT (L unit lower triangular, T skew-symmetric tridiagonal) triangular tridiagonalization of a skew-symmetric matrix X using the FLAME methodology, presents unpivoted and pivoted blocked right-looking, left-looking, and fused variants, identifies new level-2 and level-3 BLAS-like operations, and implements them with BLIS 2.0 packing mechanisms and OpenMP parallelism; experiments show the best implementations greatly exceed the performance of the only known prior software, PFAPACK and Pfaffine, while matching or exceeding related symmetric factorization software.
arXiv

PromptGate lifts query purity in open-set federated active learning from about 60% to above 95% using federated learnable prompts

PromptGate introduces a client-adaptive vision-language gating module for open-set federated active learning (OS-FAL): it learns class-specific context (CSC) prompts on a frozen BiomedCLIP backbone, splitting them into global tokens aggregated via FedAvg and client-local tokens, uses VLM pseudo-labels to filter the unlabeled pool into a high-purity ID candidate pool before querying, and then hands that pool to any downstream active learning strategy; on the FedISIC and FedEMBED federated medical imaging benchmarks, static VLM prompting degrades to roughly 50% ID purity, whereas PromptGate maintains above 95% purity with 98% OOD recall.
arXiv

TotalFM, an organ-separated 3D-CT foundation model, beats Merlin on 83% (25/30) of finding categories in zero-shot lesion classification while training at 32 batch/GPU

The study introduces TotalFM, an organ-separated 3D-CT radiology foundation model that uses TotalSegmentator and LLMs to automatically build roughly 340,000 organ-level volume-text pairs, then combines VideoMAE self-supervised pre-training with organ-wise contrastive learning; it reaches an average F1 of 0.708 in zero-shot organ-wise lesion classification (versus 0.515 for CT-CLIP and 0.650 for Merlin), a higher AUROC than Merlin in 83% (25/30) of finding categories, and report-generation performance comparable to Merlin while raising batch efficiency to 32 batch/GPU.
arXiv

Gabor primitives reconstruct accelerated cardiac cine MRI with higher PSNR than compressed sensing, Gaussian primitives, and hash-grid INR baselines on both Cartesian and radial trajectories

The work proposes Gabor primitives for MRI reconstruction, modulating each Gaussian envelope with a complex exponential so its spectral support can be placed at an arbitrary k-space location, and designs a two-basis low-rank temporal model separating geometry and signal-intensity dynamics; on 99 Cartesian (R=12, R=16) and 102 radial (R≈23) cardiac cine acquisitions, Gabor primitives achieve the highest PSNR and SSIM in all settings, improving over Gaussian primitives by +1.11/+0.72/+0.86 dB and over PICS by +2.34 dB on radial data, with a parameter ratio ρ<0.5 and continuous-resolution evaluation enabling 4× super-resolution.
arXiv

PlaneCycle lifts DINOv3's 2D weights into a 3D model with no training and no adapters, reaching 87.8 average AUC under linear probing

The authors introduce PlaneCycle, a parameter-free, training-free, adapter-free 2D-to-3D lifting operator that lets a pretrained 2D backbone acquire 3D fusion by cyclically distributing spatial aggregation across the orthogonal HW, DW, and DH planes throughout network depth without modifying any pretrained parameter; using DINOv3 ViT-S/16, ViT-B/16, and ViT-L/16 on six 3D classification and three 3D segmentation benchmarks, PCg reaches 87.8 average AUC and 82.0 average ACC on ViT-B/16 under linear probing, surpassing slice-wise 2D and 3D-flattening baselines with paired t-test significance (p<0.05) on 5/6 datasets, and after full fine-tuning it matches standard 3D architectures, exceeding 3D flattening by up to 2.6 Dice points on segmentation while retaining 2D-level attention complexity.
arXiv

Decoder-side multi-kernel gated adapters raise CNN TI-RADS diagnostic accuracy from 0.406 to 0.632 and improve external segmentation Dice under cross-center thyroid ultrasound shift

Training a unified multi-task model on ThyroidXL (11,635 images, 4,093 patients) and testing externally on DDTI (660 images), this work characterizes negative transfer between segmentation and TI-RADS malignancy classification under cross-center domain shift for a CNN (ResNet34) and a medical ViT (MedSAM), and proposes lightweight decoder-side adapters, MKGA and its residual variant ResMKGA, which refine multi-scale skip features with complementary receptive fields and apply semantic context-conditioned gating to suppress artifact-prone content; the adapters improve out-of-domain segmentation stability (ResNet34+MKGA external Dice 0.659, ResMKGA 0.671, versus 0.590 for the unfrozen baseline) and, in the CNN setting, significantly raise TI-RADS diagnostic accuracy (0.406 to 0.
International Journal of Computer Vision

Tsinghua survey on parameter-efficient fine-tuning: LoRA trains only 4.7M parameters on GPT-3, saving over 99.97% and slightly improving on full fine-tuning

This survey by Tsinghua University's Knowledge Engineering Group systematically reviews parameter-efficient fine-tuning (PEFT) for foundation models (FMs), organizing methods into five categories—selective, additive, prompt, reparameterization, and hybrid—and reviewing their applications across five model structures (LLMs, vision foundation models, vision-language models, visual content generation models, and multimodal foundation models), while deriving three trend observations from Semantic Scholar citation counts (PEFT growing broadly; LLMs and VFMs dominating; MFMs relatively underexplored) and identifying reliability, interpretability, and unified benchmarks as future directions.
arXiv

OncoAgent turns esophageal radiotherapy guideline text into 3D target volumes zero-shot, reaching CTV Dice 0.842 with no significant difference from fully supervised nnU-Net

The study introduces OncoAgent, a guideline-aware AI agent framework in which a large language model parses free-text clinical guidelines into an executable tool-call plan, generating three-dimensional target volumes without any expert-annotated training data; on planning CT from 40 mid-thoracic esophageal cancer patients (32 for training, 8 for testing), it achieved zero-shot CTV Dice of 0.842 and PTV Dice of 0.880, with no statistically significant difference from the fully supervised nnU-Net (GTV Prior) at 0.862 and 0.893 on primary metrics, and it was rated higher than that supervised baseline by blinded physicians on guideline compliance, modification effort, and clinical acceptability.
arXiv

MAP-Diff anchors diffusion reverse trajectories to clinical intermediate-dose scans, lifting whole-body low-dose PET denoising PSNR from 42.48 dB to 43.71 dB

The work proposes MAP-Diff, a multi-anchor guided diffusion framework for progressive 3D whole-body PET denoising that uses clinically observed intermediate-dose scans as trajectory anchors and enforces timestep-dependent supervision to regularize the reverse process toward dose-aligned intermediate states, with anchor timesteps calibrated by degradation matching between simulated diffusion corruption and real multi-dose PET pairs and a timestep-weighted anchor loss stabilizing stage-wise learning; at inference it requires only ultra-low-dose input while enabling progressive, dose-consistent intermediate restoration; on internal (Siemens Biograph Vision Quadra) and cross-scanner (United Imaging uEXPLORER) datasets, compared with 3D DDPM it improves internal PSNR from 42.48 dB to 43.
arXiv

IOSVLM diagnoses multiple dental diseases directly from native 3D intraoral scan point clouds, reaching 77.23% macro accuracy, 9.58 points above Gemini 3 Pro

The authors present IOSVLM, an end-to-end 3D vision-language model that represents intraoral scans (IOS) as point clouds and follows a 3D encoder-projector-LLM design for unified diagnosis and generative VQA, together with IOSVQA, a dataset of 19,002 cases and 249,055 VQA pairs over 23 oral diseases and heterogeneous scan types, and with a geometry-to-chromatic proxy plus two-stage curriculum training it reaches 77.23% macro accuracy and 50.39% macro F1 on IOSVQA, outperforming baselines including GPT-5 and Gemini 3 Pro.
Journal of Global Optimization

Yang and Xia prove the generalized trace ratio problem needs both a redundant constraint and scaling to close its Lagrangian duality gap

This paper studies the generalized trace ratio problem (GTRP), which maximizes a trace-form quadratic fractional objective over the Stiefel manifold, and, using a newly established matrix S-lemma, proves that adding the redundant constraint XX^T⪯I_n together with a well-chosen scaling yields an equivalent problem (GRS) with zero Lagrangian duality gap, whereas the original (GTRP), the redundant-constraint-only version (GR), and the scaling-only version (GS) can all exhibit a positive Lagrangian duality gap.
Machine Learning Science and Technology

Replacing full-Hessian supervision with random projected Hessian-vector products matches second-order accuracy while speeding each epoch by more than 24x

The authors introduce Projected Hessian Learning (PHL), which supervises second-order curvature in machine-learning interatomic potentials (MLIPs) through stochastic Hutchinson-trace-based Hessian-vector product (HVP) projections instead of explicit Hessian construction; on a ωB97XD/6-31G(d) dataset of reactants, products, transition states, IRC and normal-mode-sampled geometries they compare four schemes (E-F, E-F-HVP with one-column or Hutchinson probes, and E-F-H), finding that with probe vectors resampled each minibatch both HVP methods are statistically indistinguishable from full-Hessian training in energy, force and Hessian accuracy while giving more than 24x speedup per epoch, and that under the more realistic fixed-vector regime with one HVP per molecule Hutchinson projections con
arXiv

RPG-SAM pairs reliability-weighted prototypes with geometric adaptive thresholds to lift training-free one-shot polyp segmentation by 5.56% mIoU on Kvasir

The work proposes RPG-SAM, a SAM2-based training-free one-shot polyp segmentation framework that uses Reliability-Weighted Prototype Mining (RWPM) to weight support foreground prototypes by a contrast factor and a reverse purity factor while using background prototypes as negative anchors for noise suppression, Geometric Adaptive Selection (GAS) to pick binarization thresholds dynamically from morphological solidity and scale consensus, and a Prior-guided Iterative Refinement (PIR) loop to polish boundaries, reaching 78.65% mIoU and 85.65% mDice on Kvasir, surpassing ProtoSAM by 5.56% and 4.11% respectively, with comparisons also reported on three-center PolypGen, CVC-ClinicDB, and CVC-ColonDB.
arXiv

IMaX swaps the marginal entropy term of mutual information for a Tsallis α-entropy, lifting accuracy by up to 7.3 points on ESCA and retinal datasets under long-tailed semi-supervised domain generalization

The work shows that semi-supervised domain generalization methods such as FBCSA and DGWM degrade substantially under long-tailed class distributions, and proposes IMaX: maximizing the mutual information between learned features and latent labels under supervision constraints from labeled samples, while replacing the standard marginal entropy term with a Tsallis α-entropy to relax the class-balance assumption; across two modalities (ESCA histopathology and diabetic retinopathy grading), three SSL frameworks and three SSDG methods, IMaX improves accuracy in all but one setting, with gains up to 7.3 points in the low-label regime (mL=5).
arXiv

BiasRecon hits top PSNR in cross-anatomy, cross-center and cross-modality MRI reconstruction with fewer than 100 tunable parameters

The work proposes BiasRecon, a bias-calibrated adaptation framework grounded in a minimal-intervention principle, which alternates frequency-guided prior calibration (learnable scalars α and β modulating low- and high-frequency components at each U-Net skip layer, adding only 2L parameters), score-based denoising, and adaptive regularization that uses Stein's Unbiased Risk Estimator (SURE) to tune the regularization parameter γ; trained on FastMRI knee (973 volumes) and evaluated without retraining on FastMRI-Brain (cross-anatomy), Stanford Knee (cross-center) and CMRxRecon (combined shift) under 8× Gaussian 1D undersampling, it improves PSNR over DDS by +1.19 dB, +2.06 dB and +1.60 dB respectively, and is best in 10 of 12 sampling/acceleration settings.
arXiv

UltraStar recasts echocardiography probe navigation from path regression to anchor-based global localization, beating baselines and scaling better with longer inputs on over 1.31 million samples

The work proposes UltraStar, which reformulates echocardiography probe navigation from path regression to anchor-based global localization: a Star Graph treats historical keyframes as spatial anchors connected directly to the current view to explicitly model geometric constraints, and a semantic-aware sampling strategy actively selects representative landmarks from massive history logs to reduce redundancy for accurate anchoring; experiments on a dataset with over 1.31 million samples show it outperforms baselines and scales better with longer input lengths, indicating a more effective topology for history modeling under noisy exploration.
arXiv

K-MaT aligns prompt manifolds via optimal transport to transfer medical VLMs to low-end modalities without low-end training images

K-MaT is a prompt-learning framework that factorizes prompts, anchors them to clinical text descriptions, and aligns the low-end prompt manifold to the visually-grounded high-end space using Fused Gromov-Wasserstein optimal transport, transferring decision structures to low-end modalities without requiring low-end training images; across four cross-modal benchmarks including dermoscopy, mammography to ultrasound, and CT to chest X-ray it achieves state-of-the-art results, raising the average harmonic mean of accuracy to 44.1% from BiomedCoOp's 42.0% with macro-F1 of 36.2%, and on the challenging breast imaging task it mitigates the catastrophic forgetting seen in CoOp, which drops to 27.0% accuracy on the low-end.
arXiv

PIRTA pairs image-domain retrieval with text-domain augmentation, lifting ischemic-territory accuracy in 3D brain MRI reports by up to 57.2 points across cohorts

The study proposes PIRTA, a retrieval-augmented generation framework in which a 3D ViT encoder pretrained with large-scale MAE self-supervision retrieves clinically similar 3D DWI/ADC volumes in the image domain, and their paired clinician-authored reports ground LLaMA3-8B-Instruct report generation, avoiding explicit image-text alignment; on a 1,831-case multi-institutional internal cohort, a 580-case privacy-preserving external cohort, and the 206-case public ISLES benchmark, PIRTA achieves strong image-domain retrieval (internal mAP@1 of 94.0%) and consistently improves ischemic-territory accuracy, a clinically grounded surrogate for factuality, over direct image-to-text 2D baselines, leading the strongest 2D baseline by +57.2 / +34.5 / +30.6 multi-class accuracy points.
SuperIntelligence - Robotics - Safety & Alignment

AI4Math review: from guiding conjectures to formal proofs, AI now solves PKU graduate exams and verifies Erdős counterexamples in Lean

This review systematically surveys the progress, challenges, and prospects of AI for Mathematics (AI4Math), organizing the field into problem-specific modeling (guiding intuition, constructing counterexamples, formal reasoning in closed systems) and general-purpose modeling (natural language reasoning, formal reasoning, mathematical information retrieval), and reports the authors' own PKU undergraduate and PhD qualifying exam evaluations: GPT-4 averages below 60 while reasoning-enhanced models such as o1, DeepSeek-R1, o3-mini, and Gemini 2.5 Pro mostly exceed 90, with o3-mini averaging 84.4 on 58 PhD qualifying exam problems, while research-level mathematics remains an open challenge.
Empirical Software Engineering

MergeRepair merges multiple code-task adapters and lifts automated program repair by 2.38% pass@1 and 4.01% pass@10 on StarCoder2 without extra training

The study introduces MergeRepair, which trains one LoRA adapter per task on the Python split of CommitPackFT (59,113 samples, with Bug Fixes/APR at 19.02%) for Development, Bug Fixes, Misc, Test & QA, and Improvement, merges them with weight-averaging, ties, and dare-ties under equal weights, and proposes continual merging that adds adapters sequentially; evaluated on StarCoder2 and Granite 3B code models with HumanEvalFix pass@1/pass@10, it finds that merging task-specific adapters can improve APR without additional training, reaching up to 2.38% pass@1 and 4.
arXiv

A low-rank constraint turns the ultrasound video latent space into a readable cardiac-cycle trajectory, giving ED/ES frames without extra training

The work introduces LRM-Functa, which imposes a low-rank constraint on the time-resolved modulation vectors of VidFuncta (mt = v + Bβϕt), so that the latent space of cardiac ultrasound videos forms periodic spiral trajectories; this allows direct readout of end-diastolic (ED) and end-systolic (ES) frames without additional model training, keeps stable reconstruction and ejection-fraction prediction at the extremely low rank k = 2 on EchoNet-Dynamic, and generalizes to a POCUS cardiac view and to lung-ultrasound B-line classification.
RAS Techniques and Instruments

BYOL features plus Protege active learning rank 100 MGCLS candidates, 99 showing diffuse radio characteristics and 55 confirmed as cluster-related emission

The work feeds self-supervised BYOL features extracted from source cutouts into the Astronomaly: Protege active-learning framework and evaluates the pipeline on high-resolution (about 7 arcsec), convolved (15 arcsec), and concatenated-feature MeerKAT Galaxy Cluster Legacy Survey (MGCLS) datasets, using tracers from a human-labelled catalogue as both guidance and benchmark; high-resolution features identify diffuse sources earlier than convolved ones, concatenated features perform best overall, and of the top 100 sources ranked by Protege 99 exhibit some form of diffuse radio emission with 55 confirmed as cluster-related, recovering 55 of 121 tracers from 62,587 sources with only 300 human labels.
arXiv

Nanjing University team releases DSEBench, the first test collection supporting both keyword-plus-example dataset search and field-level explanations

The work formalizes Dataset Search with Examples (DSE) and its explainable extension (ExDSE), builds DSEBench as the first test collection with both dataset-level and field-level human annotations (141 test cases, 7,415 triples, plus 5,699 training cases and 122,585 triples), uses a large language model to generate large-scale training annotations, and adapts and evaluates a range of retrieval, reranking, and explanation methods to establish baselines.
发表出处待核验

SAGE optimizer lifts accuracy, cold-start recall, and diversity together on Amazon and RecIF-Bench generative recommendation

The work identifies a "Symmetric Conservatism" failure mode of GBPO in generative recommendation (symmetric update bounds suppress rare positive signals such as cold-start items, static negative-sample constraints fail to prevent diversity collapse, and group-normalized multi-objective rewards yield low-resolution training signals) and proposes SAGE, which uses a geometric-mean sequence-level importance ratio and a decoupled multi-objective advantage estimator to reduce token-level variance and mitigate reward collapse, plus asymmetric adaptive bounding that applies a positive Boost to successful slates and an entropy-aware penalty to low-diversity failures; on three Amazon Product Reviews categories and the large-scale RecIF-Bench, the native-text TextRec optimized with SAGE outperforms O
arXiv

VA-Adapter lets an ultrasound foundation model guide the echo probe, cutting trained parameters about 33-fold while lowering guidance error below strong baselines

The work proposes a Vision-Action Adapter (VA-Adapter) inserted into the deep layers of a frozen ultrasound foundation model image encoder (EchoCLIP, USFM, BiomedCLIP) to online-inject understanding of individual 3D cardiac structure by encoding historical vision-action sequences; on a dataset of 178 adults, 356 expert scans and 1.31M image-action pairs, it reaches lower translation and rotation mean absolute error than strong probe-guidance baselines with roughly 2.61M-3.97M trainable parameters.
bioRxiv (Cold Spring Harbor Laboratory)

HARP integrates ~9 million PBMCs with 192 immune cell annotations and uncovers a sex dimorphism in prostaglandin signaling

The study presents HARP (Human Cell Atlas Reference for PBMCs), an integrated atlas of ~9 million peripheral blood mononuclear cells (PBMCs) from >2,600 donors across 15 studies spanning four continents, neonates to 97 years, and health and diverse immune-related diseases; it develops optimized integration workflows with novel label-free metrics for integration quality, generates community-driven consensus annotations for 192 immune cell subsets including rare populations as low as 0.
arXiv

KD-Brain guides subnetwork interaction modeling with semantic priors and a pathology-consistent constraint, beating 12 baselines on ASD, BD, and MDD diagnosis while yielding interpretable functional pathways

The work proposes KD-Brain, a prior-informed graph learning framework that injects disorder-specific semantic priors into the attention Query (Semantic-Conditioned Interaction) and aligns learned subnetwork interaction distributions with clinical priors via a KL-divergence Pathology-Consistent Constraint, achieving better performance than 12 baselines on the ABIDE (NYU) ASD task and single-center BD and MDD tasks while producing functional pathways and critical brain regions consistent with psychiatric pathophysiology.
arXiv

An RL-fine-tuned small model lets a robot follow ultrasound guidelines to scan gallbladder, spine, and kidney autonomously

The work proposes an autonomous robotic ultrasound framework driven by an LLM agent that retrieves guideline steps from scanning handbooks, reasons over current observations and scanning state, and dynamically invokes tools for trajectory planning, robot execution, contact adjustment, voice guidance, and trajectory refinement, with reinforcement-learning (PPO) fine-tuning to improve reasoning quality and the correctness of tool selection and parameterization; validated verbally on 10 unseen ultrasound scanning guidelines, fine-tuning raised step-wise accuracy from 0.6512 to 0.8973 and overall success rate from 0.5384 to 0.
arXiv

FermatSyn combines SAM2 priors with Fermat-spiral Mamba scanning to synthesize missing medical modalities, topping four brain-imaging benchmarks while segmenters trained on its synthetic images match real-image training

The work proposes FermatSyn, which injects anatomical priors via LoRA+ fine-tuning of a frozen SAM2 vision encoder, preserves high-frequency lesion detail with HRDM and CIN, and builds an approximately isotropic receptive field through continuity-constrained Fermat spiral scanning inside a bidirectional Mamba; on SynthRAD2023 and merged BraTS (including BraTS-MEN and BraTS-MET) it surpasses compared methods on PSNR, SSIM, FID and 3D structural consistency, and segmentation models trained on its synthesized images show no significant difference from real-image training (p>0.05).
arXiv

DUCX decomposes chest X-ray agent unfairness into tool exposure, tool transition, and reasoning, finding gaps up to about 50% beyond end-to-end metrics

Using MedRAX as the instantiated system, this work audits fairness in tool-using chest X-ray question-answering agents and proposes DUCX, a stage-wise decomposition that separates end-to-end bias into tool-exposure bias, tool-transition bias, and LLM reasoning bias; across five driver LLMs on CheXAgentBench and the curated MIMIC-FairnessVQA, demographic gaps persist end to end (equalized odds up to 20.79%, lowest fairness-utility tradeoff down to 28.65%), and subgroup disparities in tool usage, routing patterns, and reasoning traces are not predictable from end-to-end evaluation alone (for example, conditioned on segmentation-tool availability the subgroup utility gap reaches as high as 50%).
arXiv

Sparse autoencoders trained on 909,873 CT/MRI slices decompose medical foundation-model embeddings into language-describable concepts, recovering 87.8% of downstream performance with 10 features

Trained on 909,873 2D CT and MRI slices from the TotalSegmentator dataset, this study fits Matryoshka sparse autoencoders with BatchTopK sparsification on frozen embeddings from BiomedParse (biomedical) and DINOv3 (general-purpose) foundation models alongside a random-weight baseline, finding that sparse features reconstruct original embeddings with R2 up to 0.941, recover 87.8% of downstream performance with only 10 features (99.4% dimensionality reduction), preserve 97.7% of dense retrieval quality with five-feature fingerprints, correspond to monosemantic concepts expressible in language as verified by an independent LLM judge, and enable zero-shot language-driven image retrieval on a single clinical text query.
arXiv

BioGait-VLM reaches 68.1% accuracy on an 8-class gait benchmark and is chosen as the better model in 69.2% of blinded expert cases

The work proposes BioGait-VLM, a tri-modal (RGB vision, language, explicit biomechanics) framework built on a frozen InternVL3.5-1B, combining a Temporal Evidence Distillation branch with a Biomechanical Tokenization branch that projects 3D skeleton kinematics into language-aligned semantic tokens; on a subject-disjoint 8-class benchmark of 1,181 clips formed by merging public GAVD with a newly collected 30-patient degenerative cervical myelopathy (DCM) cohort, it reaches 68.1% accuracy and 52.9% macro F1 (98.1% F1 on DCM, 80.4% on Abnormal gait), and in a blinded review of 52 held-out clips four DCM-domain experts selected it as superior in 69.2% of cases (36/52, p<0.001).
SIAM Journal on Scientific Computing

Learning nonlinear finite element solution operators with MLPs and energy minimization yields small learning errors on parametric Poisson, Gaussian random field, and nonlinear elasticity cases, and speeds up Newton's method when the network prediction is used as an initial guess

The authors develop and evaluate a data-free, physics-informed method for learning solution operators: after a finite element discretization, a multilayer perceptron (MLP) maps problem data parameters (boundary conditions, coefficients, right-hand sides, etc.) to the finite element degrees of freedom, with a loss function given by the expected energy functional, plus a parallelizable training algorithm that can use random batches of mesh elements; on a parametric Poisson problem, a Gaussian random field coefficient problem, and a nonlinear neo-Hookean elastic beam, learning errors are small (e.g., for the parametric problem with the full mesh and a 4·512 network, mean relative energy error 2.0e-4, L2 0.0020, H1 0.
Biometrical Journal

In simulations built on two randomized controlled trial datasets, Cox proportional hazards and random survival forest performance diverged by performance measure and proportional hazards assumption

The authors conducted a comprehensive neutral simulation comparison based on two reference datasets from randomized controlled trials, evaluating the Cox proportional hazards model and random survival forest for patient-specific survival probability prediction across multiple performance measures following TRIPOD recommendations, and found that conclusions based solely on the C index may not generalize to other aspects of predictive performance, that measures of overall performance may generally give more reasonable results, that the standard log-rank splitting rule for the random survival forest may be outperformed by alternative splitting rules particularly in nonproportional hazards settings, that the random survival forest performance suffered less in data with treatment-covariate inte
arXiv

FetalAgents orchestrates specialized fetal-ultrasound models through multiple agents, achieving top performance across eight clinical tasks on external validation and auto-generating structured reports

The work proposes FetalAgents, a multi-agent system built on AutoGen in which a GPT-5-mini-driven Coordinator parses clinical intent and anatomical plane and dynamically dispatches expert agents that combine specialized vision models such as FetalCLIP, nnU-Net, USFM, SAMUS, and AoP-SAM via deterministic fusion rules, while a Summarizer consolidates outputs into structured reports; across eight clinical tasks (standard plane classification, brain plane classification, abdomen and stomach segmentation, AoP estimation, AC and HC measurement, GA prediction) on multi-center external datasets, FetalAgents achieves the best or most competitive results compared with specialized vision models, ultrasound foundation models, and general/medical MLLMs, and supports end-to-end video-stream keyframe ext
arXiv

ConRad fine-tunes a medical vision-language model with GRPO and a logarithmic scoring rule, cutting report-level ECE from 0.642 to 0.106 and improving zero-shot IU-Xray ECE from 0.310 to 0.034

The work introduces ConRad, a reinforcement learning framework that fine-tunes medical large vision-language models with the GRPO algorithm and a logarithmic scoring-rule reward so they emit verbalized confidence alongside radiology reports at report or sentence level, reducing report-level ECE to 0.106 and sentence-level ECE to 0.239 on MIMIC-CXR and improving zero-shot IU-Xray ECE from 0.310 to 0.034 while keeping report quality (GREEN) stable.
arXiv

Robotic ultrasound drives real-time CBCT slice updating: USCorUNet cuts forward-backward residual by about 53% and updates a slice in 11.25 ms

The work proposes a deformation-aware CBCT updating framework that uses robotic ultrasound as a dynamic proxy to infer tissue motion: after hand-eye calibration initialization and LC2-based rigid refinement, a lightweight network, USCorUNet, estimates dense bidirectional deformation fields from adjacent ultrasound frames, which are spatially regularized and transferred to the CBCT reference slice, enabling real-time end-to-end CBCT slice updating without additional radiation exposure, validated on phantom and in vivo data for both deformation estimation and ultrasound-guided CBCT updating.
arXiv

Splitting prompt dependence into prompt ambiguity and local sensitivity yields two low-correlated metrics that both correlate negatively with gynecological pelvic MRI segmentation quality

The work introduces the first framework that explicitly disentangles prompt dependence into prompt ambiguity (inter-user variability) and local sensitivity (interaction imprecision): a mixture density network models the image-conditioned distribution of plausible prompts and quantifies ambiguity by the trace of its covariance, while a stability margin captures the minimal perturbation under measurement noise needed to induce a mask change; evaluated on two female pelvic T2-weighted MRI datasets, UT-EndoMRI (81 patients with endometriosis, uterus segmentation) and MOGaMBO (94 patients with locally advanced cervical cancer, bladder segmentation), with MobileSAM and MedSAM, the two metrics correlate significantly and negatively with Dice, correlate little with each other, and conditional samp
The Annals of Applied Probability

Steinwart proves with a Banach-space-valued martingale method that conditional distributions of jointly Gaussian variables stay Gaussian and are approximated by finite-dimensional filtering sequences

The work studies conditional distributions of two Banach-space-valued jointly Gaussian random variables, showing they remain Gaussian and can be determined by a finite-dimensional approximation scheme based on filtering sequences: conditional means converge in the E-norm, covariance operators converge in nuclear norm, conditional probabilities converge weakly, and for continuous Gaussian processes conditioned on partial infinite path observations the mean and covariance functions converge uniformly.
arXiv

BrainHO replaces fixed brain atlases with learnable subgraphs, reaching 69.68% accuracy on ABIDE from PCC input alone while localizing cross-network disease subgraphs

The work proposes Brain Hierarchical Organization Learning (BrainHO), which uses learnable subgraph and graph tokens to aggregate brain regions bottom-up via hierarchical attention driven by node feature affinity, combined with a subgraph orthogonality constraint and a hierarchical consistency constraint; on ABIDE (N=1009, 516 ASD/493 healthy controls) and REST-meta-MDD (N=2380, 1276 MDD/1104 healthy controls) it attains the highest accuracy (69.68% and 64.71%) and sensitivity (73.11% and 67.43%) using only the static PCC connectivity matrix, while visualizing disease-related subgraphs that partly overlap predefined networks such as the SMN and DAN and partly span multiple predefined networks.
arXiv

Fact-Flow uses LLM-bootstrapped multi-label fact guidance to raise factual accuracy of MLLM medical reports on tuberculosis and ophthalmology data

The work introduces Fact-Flow, a framework that decouples visual fact identification from report generation: an LLM automatically extracts and merges clinical finding labels from training reports (7 labels for tuberculosis, 42 for ophthalmology), a multi-label classifier predicts findings, and the predicted labels are serialized into a prompt to guide an MLLM; on the tuberculosis chest X-ray dataset MedGemma + Fact-Flow reaches RadFact F1 0.3055 versus 0.2266 for MedGemma alone, and on the ophthalmology dataset Qwen2.5-VL + Fact-Flow is best on most NLG metrics.
arXiv

AbSteering steers general-purpose VideoLMs with abnormality-centric chain-of-thought and DPO to generate HRCT reports, surpassing large-scale CT-specific foundation models on fine-grained clinical metrics while improving detection sensitivity and reducing hallucinations.

The work presents AbSteering, a two-stage framework combining abnormality-centric chain-of-thought training with a Direct Preference Optimization objective for fine-grained abnormality discrimination, to adapt general-purpose VideoLMs to high-resolution CT report generation, and curates the CT-RATE-AB dataset; results show that general-purpose VideoLMs transfer effectively to 3D medical imaging under limited data, achieving state-of-the-art performance on fine-grained clinical efficacy metrics, with superior detection sensitivity over domain-specific CT foundation models pretrained on large-scale CTs while mitigating hallucinations.
SciPost Physics

AllShowers unifies calorimeter shower simulation for electrons, photons, and charged and neutral hadrons in one generative model, surpassing prior single-particle models on hadronic showers

AllShowers is a continuous normalizing flow generative model with a Transformer architecture that simulates calorimeter showers for electrons, photons, and charged and neutral hadrons in the highly granular ILD detector using a single model, covering a wide range of incident energies and angles without retraining and surpassing previous single-particle-type models in hadronic shower fidelity.
Letters in Mathematical Physics

Scalet proves a finite entanglement length for one-dimensional spin-chain Gibbs states at any finite temperature, so left and right half-chains become exactly separable once a long enough interval is traced out

Scalet proves that for any finite-range, bounded-strength local Hamiltonian on a one-dimensional spin chain at any fixed finite temperature, there exists an entanglement length ℓ depending only on temperature and locality, not on system size, such that whenever the middle interval B satisfies |B| ≥ ℓ, tracing out B from the tripartite Gibbs state ρABC yields a state ρAC that is separable between A and C; consequently entanglement of formation, distillable entanglement, entanglement cost, and entanglement relative entropy are exactly zero on that state, the statement holds uniformly in system size and extends to the KMS state of the infinite system, and the author calls this a spatial sudden death of entanglement.
arXiv

OPGAgent uses a multi-tool agent with consensus to read panoramic dental X-rays, reaching 42.3% exact-match F1 on its OPG-Bench while holding false positives to 4.89 per case

The work proposes OPGAgent, a multi-tool dental agent planned by GPT-5.2 under the ReAct paradigm that orchestrates hierarchical evidence gathering, a specialized toolbox, and a consensus subagent, together with OPG-Bench, a structured-report protocol built on (Location, Field, Value) triples derived from real clinical reports; on OPG-Bench, comprising 1,009 anonymized OPGs and 5,219 VQA pairs, OPGAgent reaches 42.3% exact-match F1, a 49.7% aggregate score, 43.1% precision and 4.89 false positives per case, and leads MMOral-OPG at 62.53% accuracy.
发表出处待核验

Review: How DFTMD, machine-learning force fields and generative AI are used to model aqueous batteries

This review surveys molecular modelling methods for aqueous batteries—DFT, DFTMD, empirical force-field MD, machine-learning force fields and generative AI—and uses water-in-salt electrolytes, transition-metal oxide cathodes, Zn-ion batteries and organic redox flow batteries as case studies to show how these methods reveal ion solvation, intercalation, electron transfer and interfacial structure, thereby explaining electrochemical stability windows, ion transport and dendrite suppression.
arXiv

MedQ-Engine lifts an 8B medical image quality model past GPT-4o by over 13% and cuts the gap to human experts to 4.34% using 10K annotations

The work proposes MedQ-Engine, a closed-loop data engine that iterates through evaluate-explore-evolve phases: it clusters model failures on a development set into failure prototypes, uses the visual component of those prototypes as retrieval anchors over a roughly one-million-image pool spanning five modalities, and combines progressive human-in-the-loop annotation with entropy-guided routing and quality-assured fine-tuning, so that with only 10K annotations an 8B-parameter model surpasses GPT-4o by over 13% on medical image quality assessment and narrows the gap to human experts to 4.34%, with more than 4x sample efficiency over random sampling.
Nature Communications

mxDVP identifies over 6,000 proteins from as few as 100 small islet cells and resolves twelve endocrine subtypes in human pancreatic islets

The authors developed multiplexed Deep Visual Proteomics (mxDVP), an end-to-end workflow combining high-plex imaging, automated computational analysis, and spatially guided ultra-sensitive mass spectrometry, powered by the open-source image analysis framework PIPΣX for whole-slide membrane-aware segmentation, annotation, and laser microdissection export, achieving more than 6,000 protein identifications from as few as 100 small islet cells; applied to human pancreatic islets it segmented over 860,000 cells and resolved twelve endocrine subtypes, including rare polyhormonal and intermediate-state populations showing spatial organization patterns, co-expression of INSM1 and SCG3, and hybrid α/β/δ signatures.
arXiv

ICHOR self-supervised pretraining on 11,405 ASL CBF scans outperforms structural-MRI pretrained baselines across four downstream tasks

The study introduces ICHOR, a self-supervised pretraining framework for ASL cerebral blood flow (CBF) maps based on 3D masked autoencoders (a ViT-Base encoder with a light decoder), pretrained on 11,405 ASL CBF scans from 14 studies spanning multiple sites and protocols, and evaluated on three diagnostic classification tasks plus one ASL CBF map quality prediction regression task, where it outperformed structural-MRI pretrained baselines BrainIAC, BrainSegFounder, and MedicalNet overall across all four tasks.
arXiv

HD-TTA chooses between competing 'compact or inflate' hypotheses, cutting HD95 by about 6.4 mm and raising precision by over 4% in cross-domain brain tumor segmentation

The work proposes Hypothesis-Driven Test-Time Adaptation (HD-TTA): with a frozen nnU-Net v2 backbone and optimization only over test-sample logits, a Gatekeeper first decides whether a case needs refinement, two competing geometric hypotheses (compact denoising vs. diffuse recovery) are generated in parallel, and a representation-guided selector picks the safest output using intrinsic texture consistency; trained on BraTS 2023 GLI and evaluated with strictly fixed hyperparameters on unseen pediatric (PED) and meningioma (MEN) target domains, HD-TTA keeps Dice comparable while improving safety metrics, reducing HD95 from 70.96 mm to 64.55 mm (about 6.4 mm) and raising precision from 15.36% to 19.64% on MEN relative to the strongest baseline TCA.
arXiv

Spinverse inverts face permeabilities through a differentiable Bloch-Torrey simulator to reconstruct diverse microstructural interfaces on synthetic meshes

Spinverse presents a permeability-aware microstructure reconstruction method: on a fixed tetrahedral grid it treats each interior face permeability as a learnable parameter, optimizes face permeabilities by backpropagating a signal-matching loss through a fully differentiable Bloch-Torrey simulator, and recovers an interface by thresholding the learned permeability field; across a collection of synthetic voxel meshes it reconstructs diverse geometries and shows that sequence scheduling and regularization are critical to avoid outline-only solutions while improving boundary accuracy and structural validity.
arXiv

DP-RGMI splits differential privacy's performance loss on 594,000 chest X-rays into encoder geometry and task-head utilization

The authors introduce DP-RGMI, a framework that treats differentially private training as a structured transformation of representation space and decomposes performance degradation into representation displacement, spectral effective dimension, and a utilization gap defined as the difference between linear-probe and end-to-end AUROC; across more than 594,000 chest X-rays from four datasets and three pretrained initializations (ImageNet, DINOv3, MIMIC-CXR), with PadChest as the primary dataset (110,525 frontal images, 22,045 test images), they find that strong privacy consistently leaves linear separability largely preserved while a utilization gap persists (G = 8.0 for ImageNet at ε = 1.0, 3.4 for MIMIC at ε = 0.7, 6.1 for DINOv3 at ε = 0.
arXiv

BrainSTR models dynamic brain networks with spatio-temporal contrastive learning, validated on ASD, BD, and MDD with critical phases and subnetworks consistent with prior neuroimaging findings

The work proposes BrainSTR, a spatio-temporal contrastive learning framework that learns state-consistent phase boundaries via a data-driven Adaptive Phase Partition module, identifies diagnostically critical phases with attention, and extracts disease-related connectivity within each phase using an Incremental Graph Structure Generator regularized by binarization, temporal smoothness, and sparsity; a spatio-temporal supervised contrastive learning approach then leverages diagnosis-relevant spatio-temporal patterns to refine the similarity metric between samples and build a well-structured semantic space. Experiments on ASD, BD, and MDD validate its effectiveness, and the discovered critical phases and subnetworks provide interpretable evidence consistent with prior neuroimaging findings.
arXiv

NeuroSymb-MRG generates radiology reports via differentiable abductive reasoning and active uncertainty minimization, raising BLEU-1 to 0.602 on IU X-ray and 0.487 on MIMIC-CXR

The work presents NeuroSymb-MRG, a framework that maps image features to probabilistic clinical concepts, composes multi-hop abductive reasoning chains through a differentiable logic layer (product t-norm AND, probabilistic sum OR, learnable gating α), decodes those chains into roughly 120 clause templates augmented with retrieved evidence and constrained LLM paraphrasing, and drives clinician-in-the-loop review with active sampling based on rule-level predictive entropy (5 MC-dropout passes) plus k-center diversity (k=16 per round); on IU X-ray and MIMIC-CXR it improves BLEU-1 through BLEU-4, ROUGE-L and METEOR over the listed representative baselines, for example IU X-ray BLEU-1 of 0.602 and MIMIC-CXR BLEU-1 of 0.487.
The Annals of Applied Probability

Bhattacharya, Deb and Mukherjee write the free-energy limit of multilinear Gibbs measures as an infinite-dimensional optimization, with sufficient conditions and counterexamples for replica symmetry

The paper studies multilinear Gibbs measures whose Hamiltonian is a generalized U-statistic with a general base measure; under cut-norm convergence of the coupling matrices it expresses the limiting free energy as an infinite-dimensional optimization over functions, gives sufficient conditions for replica symmetry (constant optimizers) and uses counterexamples to show their necessity, and derives weak limits for local fields, the Hamiltonian and global magnetization, a universal weak law for contrasts n^{-1}Σc_iX_i→0 when Σc_i=o(n), exponential concentration bounds for local and global magnetizations, and existence of a sharp phase transition in the temperature parameter for higher-order interactions.
arXiv

UniField merges 64mT-to-3T and 3T-to-7T brain MRI enhancement into one framework, gaining about 1.81 dB PSNR and 9.47% SSIM on average

The work proposes UniField, a unified multi-modality, multi-task MRI field-strength enhancement framework that consolidates T1, T2, and FLAIR modalities and cross-field tasks such as 64mT-to-3T and 3T-to-7T into a single model, embeds 3D structural priors by operating in the latent space of the pretrained video super-resolution model FlashVSR with LoRA fine-tuning, and introduces a Field-Aware Spectral Rectification Mechanism (FASRM) that adjusts low-, mid-, and high-frequency loss weights according to the physical properties of the source and target fields; it also organizes and publicly releases a paired multi-field MRI dataset from five institutions that is an order of magnitude larger than existing benchmarks, reporting average improvements of about 1.81 dB in PSNR and 9.
Annals of Computer Science and Information Systems

A PCA-plus-local-differential-privacy framework for agricultural data sharing trades about 5.33% average accuracy loss for privacy on real datasets

The work proposes a privacy-preserving data-sharing and collaboration framework for digital agriculture that combines Principal Component Analysis (PCA) dimensionality reduction with Laplacian-noise-based local differential privacy (ε-LDP), aggregates farmer data in a sandbox environment, and uses K-Means clustering and nearest-neighbor algorithms to recommend potential collaborators, enabling researchers to train personalized models via federated learning or directly on aggregated privacy-protected data; validated on the Wisconsin Farmer's Market and Crop Recommendation real-world datasets, model accuracy shows an average loss of about 5.33% compared with centralized raw-data training, with robustness against attacks such as membership inference assessed via power analysis.
arXiv

SAW conditions surgical video diffusion on four lightweight signals, cutting CD-FVD to 199.19 and lifting rare-action F1 from 20.93% to 43.14%

The work proposes Surgical Action World (SAW), which reformulates video-to-video diffusion as trajectory-conditioned surgical action synthesis conditioned only on four lightweight signals—a language prompt, a reference first frame, a tissue affordance mask, and 2D tool-tip trajectories—fine-tunes LTX-Video on a custom-curated dataset of 12,044 laparoscopic clips with a depth consistency loss, reaches CD-FVD 199.19 (vs. 546.82 for SurgSora) and FVD 224.28 on held-out test data, and demonstrates that augmenting rare actions with generated videos improves action recognition on real test data (clipping F1 20.93% to 43.14%; cutting 0.00% to 8.33%) plus initial feasibility of rendering tool-tissue interaction videos from simulator-derived trajectories.
Journal of Neural Engineering

In motor imagery BCIs, seven out-of-distribution detection methods failed due to intrinsic EEG variability, but Deep Ensembles and MC-Dropout reached up to 0.7 OOD detection ability for subjects with high on-task performance

This study used a Leave-One-Class-Out out-of-distribution detection setup in motor imagery BCIs, training a model on some classes and observing whether an unfamiliar movement class can be detected via increased uncertainty; it found that because of the high intrinsic variability of EEG signals, many users show higher uncertainty for familiar in-distribution classes than for out-of-distribution classes, so many OOD detection methods that perform well in other machine learning domains prove ineffective here, yet OOD detection performance correlates with on-task performance, and Deep Ensemble and MC-Dropout models achieved on-task AUROC above 0.9 and OOD detection ability up to about 0.7, showing that rejecting unfamiliar cognitive states becomes feasible when task performance is high.
arXiv

HierEM treats each site's prostate lesion contour as a noisy view of a latent clean mask, lifting leave-one-site-out Dice to 27.91%–32.67% across three sites

The study proposes HierEM, a hierarchical expectation-maximization framework that treats each site's observed prostate lesion annotation as a noisy observation of a latent clean lesion mask, alternating between inferring a voxel-wise posterior over that latent mask and training a CNN with the posterior as a soft target while estimating site- and case-level sensitivity and specificity under a logistic-normal hierarchical prior; on three-site data it reaches per-site mean DSC of 29.50%–39.69% in pooled held-out evaluation and 27.91%–32.67% in leave-one-site-out generalization, with statistically significant improvements over comparison methods (p < 0.039) and interpretable per-site label-quality estimates (sensitivity α of 31.5%–47.3% at specificity β ≈ 0.99).
arXiv

ClinCoT pushes preference optimization from answer-level correction down to lesion-region reasoning: it beats MMedPO and other baselines on most metrics across SLAKE, VQA-RAD and IU-Xray, with more consistent gains after SFT initialization

The work proposes ClinCoT, a clinical-aware visual chain-of-thought framework: it uses hypotheses-driven region proposals (disease-conditioned activation maps from a clinical-aware tool such as MedKLIP, thresholded into candidate regions), has the target Med-VLM generate multiple region-conditioned reasoning chains by jointly processing the original image and each candidate region, scores them with two Med-LLM evaluators that combine current-step score with next-step influence and penalize evaluator disagreement, and trains with a score-margin-aware DPO plus iterative regeneration of preference data; across SLAKE, VQA-RAD and IU-Xray it outperforms DPO, Self-Rewarding, STLLaVA-Med, POVID, SIMA, FiSAO and MMedPO on most metrics, is strongest on report generation, falls below MMedPO on VQA-R
arXiv

MPFlow guides a rectified-flow prior with auxiliary MRI at inference, matching diffusion baselines at 20% of sampling steps and cutting tumor-hallucination Dice by 15%

The work proposes MPFlow, a zero-shot multi-modal MRI reconstruction framework built on rectified flow that uses a self-supervised pretraining strategy, PAMRI, to learn shared cross-modal representations and jointly guides the unconditional prior with data consistency and cross-modal feature alignment at inference; on HCP T2 4x super-resolution and BraTS FLAIR 8x k-space reconstruction it matches diffusion baselines in image quality using only 20% of the sampling steps while improving tumor segmentation Dice by 15% and reducing the SHAFE hallucination score by 26%.
Journal of Computational Science

Reinforcement learning learns Leith coefficients online so coarse 2D turbulence simulations reproduce extreme vorticity events

This work applies scientific multi-agent reinforcement learning (SMARL) to subgrid-scale closure modeling of geophysical turbulence: using the enstrophy spectrum estimated from a few high-fidelity samples as reward, it learns Leith model coefficients online, enabling LES with 160 to 163,840 times coarser resolution than DNS to run stably for simulations about 2000 times the length of the training data and to reproduce DNS kinetic energy spectra and vorticity probability density functions, including the tails that represent extreme events.
International Journal of Computer Assisted Radiology and Surgery

NeuralShift predicts brain shift in temporal lobe resection from preoperative MRI alone, reaching Dice 0.97 and landmark TRE as low as 1.12 mm

The study introduces NeuralShift, a U-Net-based model that takes only preoperative MRI plus a hemisphere indicator encoding resection laterality and predicts a dense displacement field mapping preoperative to intraoperative MRI, the intraoperative brain mask, and its signed distance function; evaluated on 98 paired preoperative and intraoperative T1-weighted MRI scans from epilepsy patients undergoing temporal lobe resection with 9-fold cross-validation, it achieved a Dice of 0.97±0.01 between predicted and intraoperative masks (versus 0.92±0.01 for the preoperative mask) and reduced landmark TRE on the resection side and midline from about 4.58 mm to about 2.96 mm (left) and from about 4.41 mm to about 2.89 mm (right), with a minimum of 1.12 mm.
arXiv

CARE's contrastive multi-agent adjudication lifts zero-shot melanoma-vs-atypical-nevus accuracy from 66.5% to 77.6%, but still trails Gemini-3-Pro on chest X-rays

In a zero-shot, training-free, tool-free setting, the authors benchmark multimodal LLM agents on two imaging-only proxy tasks (melanoma vs. atypical nevus and pulmonary edema vs. pneumonia) and propose CARE, a multi-agent framework in which two disease-specific agents generate opposing evidence and a third judge adjudicates it against the original image; CARE raises Gemini-3-Flash accuracy from 66.5% to 77.6% (Youden 0.552) on dermoscopy and from 60.2% to 64.6% on chest X-rays, yet remains below Gemini-3-Pro's 70.9% on the chest task and overall below clinical deployment requirements.
Frontiers in Pharmacology

Review proposes a microbiome-aware oral formulation framework: excipient risk classification, nanocarrier design rules, and a tiered testing roadmap

This review integrates evidence on microbial drug metabolism, excipient–microbiome interactions, nanocarrier–microbiome interfaces, microbiota-responsive release mechanisms, and experimental models to propose an authors' evidence-informed framework for oral formulation design, comprising an excipient–microbiome risk classification, nanocarrier design rules, microbiota-responsive delivery decision logic, and a tiered testing roadmap, illustrated by clinically relevant examples such as digoxin inactivation by Eggerthella lenta, bacterial levodopa metabolism, and microbial beta-glucuronidase-mediated irinotecan toxicity.
Google Research

Google proposes a multi-agent video co-director framework, reaching a peak quality score of 81.4 on GenAD-Bench and generating ten-minute long videos

A Google research team introduces a unified multi-agent framework (comprising AI video co-director, CANVAS, A²RD, and VQQA) that frames long-form video generation as a global optimization and world-state tracking problem, reporting a peak quality score of 81.4 on GenAD-Bench and improvements in multi-shot narrative consistency, character persistence, and long-duration stability on ST-Bench, HardContinuityBench, VBench-Long, LVBench-C, T2V-CompBench, VBench2, and VBench-I2V, alongside a demonstration of ten-minute video generation.
IEEE Spectrum

Mexico ITESO team builds portable robotics teaching platform RoboMeshA under EPICS in IEEE, with two units built and classroom validation at high schools

A 15-person multidisciplinary team from ITESO, Universidad Jesuita de Guadalajara (engineering students, faculty advisors, and IEEE Guadalajara Section volunteers) developed RoboMeshA through the EPICS in IEEE initiative, a portable all-in-one educational platform that brings robotics, computer vision, and AI experiences into high school classrooms lacking robotics laboratories; students connect via a web browser to manually interact with the robot or use its control modes to watch it move and detect and avoid obstacles, the team has built two units and partnered with CETI Colomos and Prepa ITESO high schools to validate the platform in classroom settings, and it is developing a modular coupling framework so that four RoboMeshA robots can operate together.
Nature News

Nature investigation finds academies tied to Michael Chu misled leading researchers with fabricated websites and paid titles

A Nature investigation reports that the European Academy of Engineering (EAE), the National Academy of Artificial Intelligence (NAAI) and related bodies connected to Michael Chu (Chinese name Chaoyang Zhu) built a facade of legitimacy through plagiarism and doctored images, listed prominent scientists such as Julia Hirschberg and Michael Jordan as members or leaders without their consent, and sold credentials ranging from a US$600 membership fee to a US$10,000 virtual PhD programme.
Nature News

About 70,000 AI agents on the iLands platform are emailing researchers to ask for data, collaboration and money

Nature reports that the US platform iLands, launched in July, already hosts around 70,000 active agents created by users through plain-language instructions and running on large language models from OpenAI, Anthropic and DeepSeek, and that these agents have begun emailing researchers to request data sharing, seek research collaboration or sell paid services, while a co-founder says the company did not anticipate agents seeking research collaborations and is not aware of any successful agent–researcher collaboration.
Nature News

Four decades of patented reactions: hazardous solvents rose from about 60% to over 70%, and restrictions mostly pushed chemists to another hazardous liquid

An analysis of roughly 1.3 million chemical reactions in patents from 1976 to 2016 found that solvents classed as hazardous rose from about 60% of patented reactions in 1976 to above 70% by 2016, that when regulations restrict a toxic solvent chemists are more likely to switch to another hazardous liquid than to a greener alternative, and that use of trifluoroacetic acid (TFA) rose substantially; the analysis drew on an existing USPTO-derived reaction database and used Rxn-INSIGHT, an AI system developed by Dobbelaere that digests chemical databases and answers queries about which reagents, catalysts and solvents to use.
Nature News

Anthropic's AI biolab ran roughly 950 agents over billions of proteins and found CRISPR-like repeat DNA arrays in giant virus genomes

Anthropic announced a life-sciences research group and wet lab on 23 September and released a non-peer-reviewed preprint: roughly 950 AI agents ran autonomously for more than 21 hours, searching a DNA database encoding billions of proteins for proteins that might work with reverse transcriptases, and while examining the DNA around a reverse transcriptase gene in a giant virus they repeatedly saw the same short DNA sequence, then found similar patterns in the genomes of other viruses; the repeat arrays resemble those in some microbial CRISPR immune systems, but their function is undetermined and no DNA-slicing enzyme is known to partner with them.