Skip to main content

Daily report

AI and science frontiers · 2026-09-22

Only content delivered through the publication boundary on this date is included.

Terence Tao blog RSS

Open problems, open mathematics

This guest post by Antonio Auffinger draws on his decade of conversations with biologists, computer scientists, and physicists to argue that mathematics' culture of open sharing may erode if machine proof generation becomes fast and accessible while credit systems remain unchanged, and it calls on the mathematical community to rethink incentives while embracing biology and applied sciences as sources of new mathematical questions and phenomena.
IEEE Spectrum

The Future Is Fanless: 100% Heat Capture for Liquid Cooled AI Servers

This CoolIT-provided article argues that beyond roughly 250 kW per rack a 70/30 liquid-to-air split leaves a 75 kW air load, so near-total liquid heat capture with air below 1% enables fanless AI servers; it describes heat cascading from processors into memory, networking, storage, and power components, CoolIT's modular coldplates plus conductive plates, vapor chambers, heat pipes, thermal transfer plates, and riding coldplates unified in one server loop, and its modeling that places full heat capture as standard for flagship rack-scale products through 2028.
MIT Technology Review

The Download: Why AI's Latest Breakthroughs and Fears May Be More Hype Than Reality

This edition of The Download centers on a commentary by Timnit Gebru and Emily M. Bender arguing that recent breathless claims about AI hacking, mathematical breakthroughs, and the prospect of self-improving superintelligence tend to look very different once experts examine what happened, and that there is a strong commercial incentive to overstate AI capabilities, with the illusion of speed and urgency also serving to misdirect policymakers and the public; the newsletter also rounds up the day's technology stories, including 22 nations calling for a new global AI oversight body, Texas and California moving to rein in data centers, AMD becoming the twelfth $1 trillion company, Meta's AI agent Muse topping the US App Store, an edible battery, political risk to NIH grants, epigenetic editing
NVIDIA Research

NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development

NVIDIA released Isaac ROS 5.0 at ROSCon Toronto, a collection of GPU-accelerated packages built on ROS that adds support for ROS Lyrical and Ubuntu 24.04, contributes a standard data-handling interface with the Open Source Robotics Alliance, introduces agent-ready Isaac skills and documentation, a FoundationStereo fine-tuning skill, an agent-ready FoundationPose inference library for object pose estimation and tracking with up to 5.5x faster performance, and a standalone pick-and-place skill, while showcasing ecosystem integrations such as RealSense, Intrinsic, Seeed Studio, Magna, and Ekumen across NVIDIA Jetson Orin Nano to Jetson Thor platforms for perception, navigation, and manipulation.
OpenAI

Parallel Halved Research Time and Cost with GPT‑6 Astra

Parallel's agents used GPT‑6 Astra to research and synthesize labor-market data, and according to this text the result was halving both research time and cost relative to prior models.
JB & JS open access

Generative Artificial Intelligence in Hip and Knee Arthroplasty: A Systematic Review of Emerging Clinical Applications in Patient Communication and Education, Documentation, and Decision Support

This systematic review searched PubMed and Embase (July 9, 2025) and included 23 studies to assess generative AI, mainly ChatGPT 3.5/4, in total hip and knee arthroplasty across patient communication and education (n=19), clinical documentation (n=2), and clinical decision support (n=2): blinded ratings found FAQ responses comparable to surgeon-written answers in accuracy, clarity, and completeness with better readability; consent documents showed better readability and completeness than surgeon versions; operative-report extraction reached 97.5%–100% agreement; decision support showed higher accuracy for surgical candidacy but low specificity for outcome prediction, alongside fabricated citations and limited patient trust.
Hugging Face

Transformers now runs llama.cpp quants

This Hugging Face post describes how transformers can now load GGUF quantized checkpoints directly through from_pretrained with a gguf_file argument: on Apple Silicon it reuses ggml Metal kernels distributed via the kernels library (packed-quantized weight matmul, fused normalization, flash attention, gated delta net, plus a home-grown topk for MoE routing) and trims generate overhead by dropping the redundant attention mask early and deferring the stopping check asynchronously; on a MacBook Pro M2 Max the authors measure generation speed close to llama.cpp across three checkpoints, report file-size tradeoffs for Q4_K_M, Q5_K_M and Q6_K, and state that packed inference is MPS-only with architecture coverage for Qwen3.5 dense and MoE (including compatible Qwen3.
bioRxiv

AI discovery of sequence rules for RNA polymerase II pausing revises the pause-release model of gene activation

Using DEFT, an LLM-guided interpretable decision-tree framework, the authors found that pause sites are discriminated by a G at the pause position, C or T at +1 and a minimum G content in the 49 upstream bases (held-out AUROC 0.92, accuracy 0.87); inserting a 377 bp G-less cassette at the 5' ends of NDRG1 and HSP90AA1 abolished the promoter-proximal pause in an orientation-dependent way without impairing hypoxic or heat-shock activation (NDRG1 was even super-activated), while ChIP 5' RNAPII, Ser5P and NELF peaks persisted without NET-seq-detectable pausing and CTD deletion weakened the pause and shifted its peak from ~+80 to ~+40, supporting a model in which the pause is initiated by G-dependent arrest and stabilized by factors including CTD-mediated tethering.
发表出处待核验

Literary History of Science: From Reassessing Darwin to Stapledon's Evolutionary Science Fiction

This chapter argues that the literary historiography of science (LH) offers its own fresh perspectives on the history of science, using the history of evolutionary theory as its case: it shows how LH scholarship and methods have been integral to the re-evaluation of Charles Darwin's significance by both historians of science and literary scholars, and it models a history of evolutionary theory extending well beyond the conventional "Modern Synthesis" of Darwinian natural selection and Mendelian genetics by showing how Olaf Stapledon's evolutionary science fiction has shaped disciplines from astrobiology and Earth system science to the theorisation of Artificial Intelligence.
bioRxiv

Expansion of DNA-Encoded Library Hits Using Generative Chemistry and Ultra-Large Compound Catalogs

This work initialized and biased the HIDDEN GEM structure-guided generative virtual screening workflow with screening data from a focused DNA-encoded library against the 53BP1 tandem Tudor domain (UNCDEL003, 58,080 compounds), nominated 57 purchasable compounds from the roughly 37-billion-compound Enamine REAL Space, and validated 14 as active hits by TR-FRET displacement (3 with IC50 ≤50 µM and 11 with IC50 ≤100 µM), with the AI-nominated hits showing greater chemical diversity, improved drug-likeness, and off-the-shelf purchasability relative to the initial DEL hits.
Journal of Artificial Intelligence General science (JAIGS) ISSN 3006-4023

Performance, Attention, and Authority: Rethinking AI-Enabled Enterprise Performance Management

This conceptual article argues that although enterprise performance management has moved from static periodic reporting to continuous real-time monitoring, finance leaders still obtain abundant data without faster or firmer decisions, locating the binding constraint in scarce managerial attention and ambiguous decision authority, and proposing a three-layer architecture of performance, attention, and authority, a five-dimension Management Attention Test (materiality, persistence, confidence, decision ability, ownership), a separation of analytical from decision authority, and governance for responsible AI use in finance, concluding that disciplined filtering and accountability serve organizations better than more AI dashboards, with AI as an analytical partner rather than a decision substi
Scientific Reports

Deep learning quantification of mouse nesting behavior for tracking cognitive decline in models of aging & Alzheimer’s disease

The study developed an AI-powered image analysis pipeline, “CozyScores,” built on the ResNet18 architecture, to score progressive nest organization on a 5-point system in young and old wild-type and 3xTg-AD mice from 0.5 to 24 h, and it found a significant age- and genotype-dependent difference at the 4 h benchmark (Young WT > Young AD > Old WT > Old AD, P < 0.05), lower maximum 24 h performance in Old versus Young groups, and generally better performance in males than females.
Blood Advances

AI-derived whole-body MRI metrics in multiple myeloma: treatment-related body composition change and its association with outcomes

This study retrained a T1-weighted Dixon whole-body MRI deep-learning segmentation pipeline, originally developed on healthy UK Biobank participants, on scans from patients with multiple myeloma to automatically generate 11 image-derived phenotypes of non-diseased tissue (volumes of abdominal subcutaneous adipose tissue, visceral adipose tissue, abdominal skeletal muscle, liver, spleen, both kidneys, both iliopsoas muscles and heart, plus liver relative fat fraction), measured baseline and longitudinal values in a 69-patient prospective observational cohort (iTIMM) undergoing induction therapy and autologous stem cell transplant, and explored associations with progression-free survival, reporting a mean Dice of 0.916 and mean Likert of 4.
Applied Intelligence

Context-Aware Semantic Similarity Measurement for Unsupervised Word Sense Disambiguation

The paper proposes a context-aware semantic similarity (CASS) method for unsupervised word sense disambiguation: given a target word, its context and an exclusion list, the context and the context with each candidate synonym substituted are encoded as vectors, and the synonym that minimizes the resulting semantic change (maximizes cosine similarity) is chosen as the sense decision; evaluated for accuracy on CoarseWSD-20, a Wikipedia-derived, noun-only benchmark of 20 words with 2 to 5 senses and 10,196 instances, against a random-option baseline (43.73%) and a most-frequent-sense baseline (73.43%), the BERT-embedding configuration reached 77.74% (7,927 hits), the only setting above the strong baseline, while USE (71.94%), ELMo (68.75%) and WMD (60.
medRxiv

Uncoded Clinical Features from Multilingual Electronic Health Records in Catalonia: Development and Validation Study

This study deployed the 3.8B-parameter open-weight small language model Phi4-mini (via Ollama) within an institutional firewall, combined with deterministic regular expression post-processing, to extract six uncoded urinary tract infection clinical features (fever, nitrites, leukocytes, lumbar pain, abdominal pain, and haematuria) from 15,498 Catalan/Spanish bilingual MEAP primary care narratives in the SIDIAP database in Catalonia, achieving 93.8% accuracy, 96.6% specificity, and 83.6% sensitivity in a double-blind clinician gold standard validation of 60 real patient records, and 87.2% accuracy, 99.5% specificity, and 99.2% positive predictive value in adversarial synthetic stress-testing of 720 notes, with no patient data leaving institutional servers.
bioRxiv

A Graph-based QSAR Modeling Pipeline for Predicting In vitro PubChem Assays and In vivo Human Hepatotoxicity: Mechanistic Analysis of Caspase-3/7 Activation

This study developed a graph-based QSAR modeling pipeline integrating assay data preprocessing, fingerprint and molecular graph feature representations, and benchmarking of classical machine learning, graph neural networks, graph transformers, and their consensus ensembles, applied to predict Caspase-3/7 activation, mitochondrial membrane potential disruption, and FDA drug-induced liver injury, where Graphormer achieved the highest F1 of 0.79 and the full consensus model achieved the highest AUC of 0.69 on DILI prediction, surpassing the previous best model with AUC 0.63 and F1 0.65, and identified structural motifs associated with dual activation and cell-line-specific responses through fragment enrichment analysis.
Health Care Management Science

Counterfactual Prescriptions via Hierarchical ML for Missed Chemotherapy Appointment Prevention

Using 1,825,948 chemotherapy appointment records from the Dana-Farber Cancer Institute, this study builds a hierarchical machine learning pipeline that first predicts cancellations and then no-shows, achieving F1-scores of 0.76 and 0.82 and improving minority-class performance by 7–10 points over a single-stage multinomial baseline; it further uses semi-supervised learning to infer no-show reasons from short-notice cancellations (weighted F1 of 0.57 and 0.54) and applies counterfactual simulation to evaluate interventions, finding that standard reminders are less effective than previously reported while provider consistency and commitment-based scheduling can reduce cancellations and no-shows.
bioRxiv

Passenger co-deletion confounds glutaminolysis signatures anchored on PTEN loss: a cautionary case for location-aware signature design

Using GISTIC copy number from the TCGA PanCancer Atlas to classify tumors as PTEN intact, hemizygous, or homozygous deletion, this study scored a five-gene glutaminolysis signature (GLS, SLC1A5, GOT1, GLUD1, GPT2) against loss severity across fourteen tumor types and found that the signature decreased with PTEN loss in all fourteen (significantly in twelve) but was not MYC-mediated; instead the decline tracked chromosomal position, since GLUD1 and GOT1 flank PTEN on 10q and thirty-seven neighboring genes carrying no glutaminolysis annotation tracked PTEN copy number just as closely (mean rho 0.843 versus 0.842), with co-deletion fidelity falling monotonically with distance from PTEN (rho = -0.
Nursing education perspectives

Harnessing Generative Artificial Intelligence and a Persona Prompt to Develop Competence in Addressing Social Determinants of Health

A family nurse practitioner (FNP) program had 98 students use a generative AI persona prompt in Claude Sonnet to create a virtual patient on a video-conferencing platform, and after a 30-minute prebriefing, about 75 minutes of scenario work, and a 45-minute PEARLS-based debriefing, the most common one-word reaction was “helpful” (n = 83), 95% said the activity facilitated their learning about SDOH (n = 82), 78% chose the two highest comfort levels for screening for and addressing SDOH (average 4.1/5), the SDOH quiz average was 94% (n = 98), and 54% answered the question on specific strategies to address SDOH correctly.
medRxiv

Benchmarking open-source automated thigh muscle MRI segmentation algorithms

This preprint benchmarks eight open-source thigh muscle MRI segmentation tools on the MyoSegmenTUM, AIPS and Sheffield base datasets plus derived pathological and augmented out-of-distribution sets, combining Dice, Jaccard, Hausdorff, boundary IoU and inter-slice Dice ratio with qualitative usability review, and finds that domain-specific U-Net models, especially MuscleMap, generally outperformed foundation-model and newer general-purpose approaches, that adding SAM variants usually degraded rather than improved segmentation quality, and that most tools showed reduced accuracy on pathological cases.
medRxiv

Soft Temporal Scoring Using a Foundation Model: Optimal Frame Selection for Improved ONSD Measurement in Ultrasound Videos

The study presents a sparsely supervised AI framework in which frozen ultrasound foundation model (USFM, ViT-B) embeddings of 768 dimensions per frame are passed to a lightweight bidirectional LSTM temporal head that outputs a 0–1 frame-quality score, trained with Gaussian soft labels that peak at expert-marked key frames and decay smoothly with frame distance (width σ = 5 frames); in subject-level five-fold cross-validation on 18 subjects and 323 ultrasound videos spanning nine controlled acquisition sweep types (about 45,500 frames), it selected a usable frame in 82.2% of sweeps containing key frames with a mean minimum distance of 3.07 frames, exceeding the strongest training-free baseline (USFM feature cosine similarity, 49.2%) and a hard-label model (71.9%).
medRxiv

Clinical trajectories and genetic architecture across the neurological–psychiatric boundary

The study compared disease-trajectory embeddings from Delphi-2M, a transformer trained only on the health records of about 400,000 UK Biobank participants, with genome-wide genetic correlations for 19 neurological and psychiatric disorders, finding moderate convergence across the 171 disorder pairs (Mantel r = 0.33, p < 1×10⁻⁴), with both measures keeping sixteen disorders closer to their own diagnostic category and making the same three exceptions — multiple sclerosis, migraine, and essential tremor sat closer on average to psychiatric disorders — so clinical trajectories and genetic architecture draw the same boundary and break it in the same places.
Amazon Science

Amazon and Stanford Launch Joint Research Initiative to Advance AI and Science

Amazon announced the Stanford and Amazon Research Initiative with Stanford University, a formal framework to advance research at the frontiers of AI, energy, and healthcare through joint research projects, PhD fellowships, and symposia, aiming to move breakthrough research into real-world solutions faster and broaden participation from diverse scholars.
NVIDIA Research

NVIDIA Launches DSX Ready: A Qualification Program for Power and Cooling Products in AI Factories

NVIDIA introduced DSX Ready, a qualification program for partner products and solutions that meet applicable NVIDIA DSX AI factory reference design requirements, launching with two initial categories, battery energy storage systems (BESS) and cooling distribution units (CDUs), and naming the first qualified BESS and CDU suppliers.
Amazon Science

Designing and Characterizing Antibodies with AI: Three Efforts Spanning Epitope Selection, Affinity Ranking, and Developability Prediction

This article presents three efforts from Amazon Bio Discovery: MochiBind, a sequence-only predictor that reframes binding affinity as pairwise comparison aggregated by TrueSkill into a global ranking, achieving higher pairwise accuracy than every structure-based baseline on four held-out antigens and scoring 200,000 antibody pairs in roughly 13 seconds on a CPU; CA-MAP, a context-aware multi-property predictor that uses example antibodies in the prompt to absorb batch offsets, holding a 0.99 correlation under a simulated batch effect where standard fine-tuning falls to 0.
NVIDIA Research

Why Deploying Physical AI at Scale Demands Safety at Every Layer

This article explains why physical AI—autonomous vehicles, humanoid robots, and industrial robots—needs safety spanning hardware, software, AI behavior, operating environment, and the deployment lifecycle as it moves from research to large-scale deployment; it presents NVIDIA Halos as a full-stack safety system across AV and robotics lines (including DRIVE AGX Thor, Hyperion, Halos OS, Alpamayo, IGX Thor, Holoscan Sensor Bridge, Isaac Lab, and Omniverse), and lists the automakers, robotics firms, chip and sensor suppliers, and certification bodies in that ecosystem, along with third-party assessments and accreditations by TÜV SÜD, TÜV Rheinland, and ANAB.
NVIDIA Research

From Enablement to Execution: Egypt and Africa's AI Ecosystem Reaches Production Scale

Using a reception at the Grand Egyptian Museum as its entry point, this article traces how Egypt's and Africa's AI ecosystems are moving from a "potential" narrative to production-scale execution, citing a more than tenfold one-year growth in Egypt's Deep Learning Institute learner base, an estimated $400 million Hassan Allam and A15 data center investment, four AI factories announced or online across Africa within a year with another 656 megawatts in the pipeline, Cassava Technologies' deployments across Egypt, Kenya, Nigeria and Morocco that could reach $720 million, Stratos Lab's more than 50 NVIDIA HGX B300 systems delivering over 7 exaflops, Morocco's Nexus AI factory outside Casablanca with a $1.
arXiv

Mira-Scene replaces sparse pose regression with pixel-aligned canonical coordinate maps, lifting 3D-IoU from 0.520 to 0.727 on BlendSwap

Mira-Scene is a compositional single-image 3D scene reconstruction framework that pairs a pixel-aligned Canonical Coordinate Map (CCM) with a scene-space Point Cloud Map (PCM) from monocular geometry estimation to form dense, bounded correspondences, recovers object transformations through robust geometric alignment, and jointly generates object geometry and CCMs with a multimodal diffusion transformer; across indoor, outdoor, synthetic, and in-the-wild scenes it improves layout accuracy over strong baselines, achieving relative gains of 39.8% in 3D-IoU and 16.5% in 2D-IoU over SAM3D on BlendSwap while using limited open-source training data.
arXiv

EdgeGen extracts compliance rules from agent policy documents and generates database-grounded rule-violating tasks, lifting finetuning mean progress by 2% to 42% on the tau2-bench airline domain and improving Gemma-4-e4b harness optimization by 30% over the base harness

EdgeGen automatically extracts compliance rules from an agent's policy document, enumerates combinations of rule violations, and grounds each scenario in executable database states via a scenario-aware SQL agent, admitting tasks only after data-grounding and scenario-feasibility checks; combined with a synthetic database generation method it forms a fully automated closed loop requiring no human annotation, yielding a 2% to 42% mean progress improvement from finetuning on the tau2-bench airline domain while some baselines degrade for some models, and harness optimization improving 10% and 30% over human-curated and base harnesses for Gemma-4-e4b.
arXiv

onPanda cuts median annotation time by 51.5% with locate-correct-continue token-level annotation, and releases the Panda-CVL dataset and benchmark

The work presents onPanda: while reading a model response, the annotator locates the first inappropriate token, either picks a substitute from the model's candidate tokens or types the correct text, and the system truncates everything after that position and continues generation from the corrected prefix, repeating this loop; a small controlled study suggests a 51.5% reduction in median annotation time over manual post-editing, 97.0% of tokens in the resulting data are model-generated, and the data carries position-precise, naturally paired positive-negative token-level correction supervision, alongside the released Chinese vision-language Panda-CVL dataset and a token-level correction benchmark.
arXiv

Bringing survey small-area estimation into AI evaluation: PP-S and PP-TS improve point and interval estimates on a benchmark and on deployed traffic, while DB-CV picks as well as an independent validation sample from one sample

The work frames disaggregated AI evaluation as finite-population survey sampling, proposes prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's prediction-powered estimate, plus a taxonomy extension (PP-TS) that borrows strength along a nested reporting hierarchy, and derives an approximately unbiased design-based cross-validation score (DB-CV); on the Open LLM Leaderboard (9,324 questions, 34 task types) and PRISM deployed traffic (68,371 rated responses, 21 LLMs by 3 conversation types), where every outcome is observed, the smoothed estimators beat direct estimators in point and interval estimation with near-nominal 95% coverage, and DB-CV selects as well as an independent validation sample at the same budget while estimating the chosen estimator's error far more ac
arXiv

D-RAC turns 236 PDFs into 1,748 retrieval-ready chunks with one multimodal pass, cutting chunking cost 77.8%–85.6% below frontier-model agentic chunking

The work presents D-RAC, which deterministically normalizes any enterprise document to PDF, applies a single multimodal LLM pass to convert rendered pages into retrieval-optimized Markdown, and then reuses W-RAC's ID-level chunk planning; on the 236-document, 795-page PDF subset of RAG-Multi-Corpus it converted and chunked the whole corpus in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks, reducing chunking-stage output tokens by 95.7%, cutting chunking cost by 77.8% (GPT-4.1) to 85.6% (Gemini 2.5 Pro), reducing chunking time by 75%, and reaching Recall@6 0.798, MRR 0.690, and NDCG@6 0.801 over 762 queries.
arXiv

Reading Without Unrolling: Virtual Unwrapping and Lead-Ink Imaging of Herculaneum Papyri

Together, an external news story and two papers present two independent routes: one tests whether imaging lead rather than carbon in the ink can make text legible, using home-made burnt papyri, while the other uses synchrotron phase-contrast micro-CT plus machine learning to achieve the complete virtual unwrapping and reading of an intact, unopened Herculaneum roll, PHerc. 1667, recovers title and author evidence identifying PHerc. 139 as Philodemus, On Gods, Book 8, and provides volumetric ink validation in PHerc. Paris 4.
Nature News

Nature maps the AI-extinction debate: an Anthropic researcher's resignation post drew over 100 million views, while a RAND assessment finds nuclear extinction infeasible but biological and atmospheric scenarios not ruled out

This Nature explainer traces the September 2026 AI-extinction debate triggered by Anthropic researcher Jacob Coxon's resignation, in which Coxon said "the people building AI earnestly believe that it could kill us all by the end of the decade," Evan Hubinger estimated the risk of human extinction at ">10% within the next decade," and Dario Amodei then posted an essay calling for a slowdown rather than a halt in AI development; it also cites RAND's Michael Vermeer, whose assessment of nuclear weapons, biotechnology and deliberate atmospheric modification found complete extinction by nuclear weapons infeasible while the other two scenarios could not be ruled out, though carrying them out would require models to have considerable ability to physically interact with the world and would almost
arXiv

Making KDA's gate signed lets a single CKDA layer track finite groups and extrapolate periodic waveforms, with 1.3B downstream accuracy on par with KDA

The work introduces Complex KDA (CKDA): extending Kimi Delta Attention by allowing signed gate entries and an extended delta-rule coefficient range so that a single diagonal-plus-rank-one transition can realize a 2D rotation; the authors prove every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition, show one CKDA layer tracks every finite group isomorphic to a subgroup of O(2), save one layer relative to comparable diagonal-plus-rank-one linear RNNs, and report experiments on group-word problems, periodic audio continuation, and language modeling.
arXiv

CARE turns robot execution failures into training data, lifting average task success by 14.5 points in simulation and 15.9 points in the real world

The work proposes CARE, which collects failed rollouts, models stage-conditioned post-failure geometric deviation distributions, synthesizes representative failure states and corrective demonstrations from them, and at inference combines stage-wise planning with physically grounded 3D point-cloud monitoring to trigger atomic adjustments or re-operations; across RoboTwin 2.0, RoboFactory, the newly introduced FSR-Bench, and real-world dual-arm tasks it reports average task-success gains of 14.5 points in simulation, 15.9 points in the real world, and a 7.5-point gain in recovery success on FSR-Bench.
arXiv

PIR reads the answer a model internally recognizes, separating "won't say" from "doesn't know"

The authors adapt the forensic Concealed Information Test to language model internals as Probe of Internal Recognition (PIR), a reference-free readout that needs no honest reference model and no labeled truth corpus: given a question and its candidate answers, it reads from the model's internal states which candidate the model recognizes as correct; across eight models from five families (Gemma, Qwen, Llama, Mistral, Phi), PIR recovers the concealed answer at 0.70 to 0.87 balanced accuracy, well above the 0.28 to 0.40 unknown-item baseline and the 0.25 chance rate, stays at 0.85 to 0.93 recognition under prompted deception, trained sandbagging, and external password-locked and circuit-broken checkpoints, and falls to the unknown-item baseline when unlearning actually removes the knowledge.
arXiv

GAM swaps language and video pretraining for a frozen 3D grounding backbone, lifting RoboTwin 2.0 randomized-scene success from 30.4% to 47.6%

The work proposes Grounded Action Model (GAM), which builds a robot foundation model on a pretrained promptable 3D grounding model (WildDet3D), maps language, 2D point, or 2D box prompts into a shared object-centric representation (image tokens keeping target visual features plus detection tokens encoding point clouds, metric positions, and extents), and uses a multi-stream transformer (MM-DiT) with robot state history to predict action chunks; it reaches 55.3% average success across 50 RoboTwin 2.0 tasks (vs. 52.0% for Spatial Forcing), 47.6% under scene randomization (vs. 30.4% for Abot-M0), a state-of-the-art 61% average across 16 LIBERO-PRO perturbation settings (vs. 53%), 17/20 successes under visual shift on a bimanual YAM (vs. 4/20), and 64.7% in-distribution and 49.
arXiv

ShieldVLA gates VLA fine-tuning with Hamilton-Jacobi reachability, cutting cumulative safety cost by 57% and lifting success rate by 0.13

ShieldVLA is a safety-aligned fine-tuning framework for vision-language-action (VLA) models that learns a model-free Hamilton-Jacobi reachability value function from visual observations as a safety critic, uses that critic to gate policy optimization into reward maximization inside feasible regions and recovery near unsafe states, and replaces manual per-step cost labels with rubric-based VLM safety scores; across five navigation and manipulation benchmarks (Dubins-VL, TurtleBot-Nav, Safety-CHORES Nav/Fetch, Franka-Reach) and multiple VLA backbones it reduces cumulative safety cost by 57% on average and improves task success rate by 0.13 over SafeVLA.
arXiv

ByteDance Seed and UW–Madison make full-pipeline FP8 RL work: tracing entropy surges to over-clipping and using Calibrated Clipping to match BF16 from 8B to 32B with up to 1.5x training throughput

This work studies the training stability of full-pipeline FP8 reinforcement learning for LLMs, identifies that compounded FP8 quantization noise distorts the importance ratio so that negative-advantage tokens are over-clipped at the lower bound and their gradients are erroneously zeroed, and proposes Calibrated Clipping, which dynamically realigns clipping bounds by matching the lower-bound clipping quantile of a BF16 reference and rebalancing the upper bound, eliminating entropy surges and restoring near-BF16 performance across GRPO and DAPO, 8B to 32B models, and multiple FP8 scaling granularities, while reaching up to 1.5x the BF16 training throughput.
arXiv

RoboDawn lets a frozen VLM control robots zero-shot through discrete commands, with one demonstration lifting RoboTwin 2.0 C2R success from 53.2% to 73.6%

The work introduces RoboDawn, which exposes robotic control to a frozen VLM as a compact set of discrete translation, rotation, and gripper commands in a closed observe-reason-act-reobserve loop, paired with an in-context learning scheme; it reaches 53.2% zero-shot and 73.6% one-shot success on RoboTwin 2.0 C2R, improves RoboDojo from 35.67% to 47.17%, and transfers zero-shot to a Franka robot for block-in-basket and block stacking.
arXiv

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

The work presents WorldCrafter, a camera-controllable autoregressive video world model that learns a camera-queryable implicit 3D-aware memory, compressing historical latent frames into a fixed budget of memory tokens read out under the requested viewpoint and injected into the video diffusion transformer before denoising, so that a single image or text prompt supports streaming, minute-scale scene exploration with better revisit consistency and camera-control accuracy than the evaluated baselines while preserving visual quality.
arXiv

Splitting repository-level SWE tasks into category experts and distilling them back into one model lifts Pro-618 mean resolution to 58.04%

Addressing a "category see-saw" in repository-level software engineering, where pooled agentic reinforcement learning improves some task categories while regressing others and aggregate resolution hides the change, the work builds a category-aware expert-training and policy-integration framework: SWE Labeler groups tasks by repository domain into service/data/security (A), user-facing applications (B), and systems/tooling/runtimes (C); Agentic-miniRL with a Refresh–Repair–Expand loop trains same-origin category experts; and label-routed multi-teacher on-policy distillation (MOPD) consolidates them into one deployable student, which reaches 58.04% mean resolution on Pro-618 (+5.39 points over base) and 59.00% on SWE-bench Multilingual (+2.
arXiv

UltraTex pushes multi-view diffusion to 2048 resolution with background token dropping and block-sparse attention, reaching up to 91.1x training speedup

UltraTex is an end-to-end multi-view diffusion framework that scales image-guided 3D texturing from 512/768 to 2048 resolution through Background Token Dropping, Block-Sparse Attention, and Foreground-Aware VAE Decoding, and it builds G-buffer TexVerse covering over 268,000 3D assets, achieving 20.6-91.1x training speedup and 22.3-74.6x end-to-end inference speedup within the dataset's common foreground-ratio range of 10%-30%.
arXiv

SkillSpec turns skill correctness into intent-masked specification reasoning: 763 confirmed defects across 515 real-world skills, 67.1% precision on code nodes

SkillSpec is a Hoare-style specification-reasoning framework that converts heterogeneous skill repositories into a unified graph, derives ExpectSpec and FactSpec under an intent mask with holistic, lineage, neighbor, and self visibility layers, and validates candidate defects in an isolated sandbox; on 515 real-world skills from SkillsBench and widely downloaded repositories it confirmed 239 skills containing 763 manually confirmed defects, with 61.2% overall precision, 67.1% on code defects and 55.4% on workflow defects.
arXiv

Harness-Zero distills an evolved agent harness into Qwen3.5-9B weights, lifting macro-average success from 23.3% to 44.3% and beating the 41.7% of the base model with the harness still attached

The work introduces Harness-Zero: an agent-as-harness setup in which a harnessing agent reviews and minimally rewrites the student agent's proposals at its response boundary, translating guidance from an evolved harness (tools, middleware, skills, memory) into executable supervision inside the target harness's action space, followed by SFT on the reviewed trajectories; across SpreadsheetBench Verified, AppWorld, and USPTO retrosynthesis, training-free agent-as-harness averages 81.1% over six benchmark-model settings versus 78.1% for code-as-harness, while the distilled Qwen3.5-9B under a minimal target harness alone raises the macro average from 23.3% to 44.3%, surpassing the 41.7% of the base model with the evolved harness still attached, and recovers 82.
Hugging Face Daily Papers

GameHorizon Suite Brings Multi-Horizon Data and Evaluation to Gameplay Research

The evidence bundle presents a Hugging Face paper page entry titled GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay, pointing to multi-horizon data and evaluation in gameplay, but neither the page body nor any referenced papers were included, so no concrete method or result can be reported.
arXiv

Functionalizer splits casing, diacritics and repetition into reversible opcodes, cutting vocabulary needs by up to 19.7% under corpus exhaustion and lifting a 98M-parameter GPT-2's Python syntax validity from 7.70% to 9.12%

The work presents the Functionalizer, a lossless pre-tokenizer framework that factors orthographic variation such as casing, diacritics and character repetition into a compositional opcode/operand prefix stream encoded in the Unicode Private Use Area, achieving complete corpus coverage with up to 19.7% fewer required vocabulary slots and, on 98M-parameter GPT-2 models, raising Python syntax validity from 7.70% to 9.12% while reducing duplicate n-gram repetition on FineWeb-Edu prose from 66.0% to 55.8%.
arXiv

ACLArena quantifies catastrophic forgetting across a four-stage post-training pipeline and proposes a LoRA expert mixture that lifts AIME26 from 10.21 toward independently trained experts

The work builds ACLArena, a four-stage sequential post-training pipeline spanning math, search, e-commerce, and instruction following; diagnoses forgetting and transfer at both the model level and the token level; systematically compares multi-teacher on-policy distillation (MMOPD), self-distilled fine-tuning (SDFT), and model merging; and proposes Mixture of Low-Rank Experts (MLE), which first consolidates multi-domain trajectories into a shared backbone via SDFT, then freezes the backbone and trains one RL-optimized LoRA expert per stage routed by environment context, improving both in-domain and out-of-domain performance on four reasoning and agentic tasks.
arXiv

1% of Tokens Can Match Full Distillation: IER Re-ranks Sparse Supervision by Gradient-Estimation Reliability

The work reframes token selection in on-policy distillation (OPD) as a gradient-estimation reliability problem at a fixed prefix, decomposes the one-sample reverse-KL gradient into signal and sampling noise in information geometry, proposes the information-efficiency ratio (IER) as a signal-to-noise ratio under the optimal scalar baseline, and approximates it on a top-K candidate set for ranking; on mathematical and medical reasoning, IER alone approaches or exceeds full OPD at token budgets of 0.1%–1%, and combined with existing usefulness scores (IER-OR, IER-AND) it matches or exceeds full OPD in multiple settings.
Nature Human Behaviour

Loneliness and desire for thinness as pathways to girls' higher internalizing problems in the Tokyo Teen Cohort

In the Tokyo Teen Cohort, 3,171 adolescents (46.9% girls) assessed at ages 12, 14 and 16 showed higher internalizing problems among girls at age 16 — depressive symptoms 15.9% versus 8.2% and suicidal ideation 29.8% versus 18.1% — and, after coproduction with young people identified loneliness and desire for thinness as candidate pathways, mediational g-formula estimates attributed 34.7% of the total effect of female sex on depressive symptoms and 46.3% of that on suicidal ideation to pathways through these factors at age 14, suggesting that social intermediates may contribute to the higher prevalence among girls in this cohort.
arXiv

Scoring the scorer: mutation analysis finds the official GPU-kernel check misses one in six witnessed faults, with 78.6% of precision faults escaping

The work adapts mutation analysis into an adequacy metric for graded numerical oracles, injecting 10,303 compilable faults into verified CUDA implementations of 188 KernelBench problems (7,384 with independent kill witnesses), measuring that the official check deterministically misses 16.9% of witnessed faults with misses skewed by family (8.7% arithmetic, 78.6% precision), and using that measurement to explain mechanisms, audit existing patches, and synthesize two-input suites reaching 98.0% detection.
arXiv

HuRo turns five human-video sources into about 142 million robotized frames, lifting real-robot OOD completion from 34.9% to 72.2% in ALLEX VLA pretraining

HuRo introduces a robotization pipeline that converts heterogeneous egocentric human videos into robot-aligned observations and action trajectories, builds the HuRo dataset of about 142 million frames from five sources (EgoDex, EgoVerse, Ego4D, Ego10K, and EPIC-Kitchens), and uses it to pretrain VLA policies on the ALLEX bimanual dexterous robot: as robotized data increases, average completion across four real-world manipulation tasks rises from 51.5% without pretraining to 80.3% with the full dataset, with ID completion from 68.1% to 88.4% and OOD completion from 34.9% to 72.2%.
arXiv

Regularizing the Self-Improvement of Agent Harnesses: How RRSI Makes Recursive Improvement Transfer

The work introduces RRSI, which keeps an agent harness (prompts, control flow, tooling, memory and context management) fully editable while regularizing both the proposal and the selection of edits, gaining up to 14.1 points on the split it evolves against and up to 4.7 points on five out-of-distribution benchmarks across eight benchmarks in coding, agentic workspace and engineering design, while running on 30% fewer policy tokens than unregularized evolution.
arXiv

Deep Persona uses a three-layer persona architecture and an ADOS-inspired evaluation framework, making two clinical simulation personas statistically indistinguishable from human dialogue under a combined baseline

The work introduces Deep Persona, a psychologically grounded three-layered architecture that organizes personas into observable expression, latent beliefs, and core motivational drives, governed by scripted determinism and bounded agency, together with a reference-free evaluation framework (ADOS-inspired metrics for pragmatic fluency, joint attention, affective congruence, and emotional expression diversity, plus a Mahalanobis-distance Dialogue Naturalness Score, DNS); on human-human and human-LLM dialogue datasets such as DailyDialog and CounselChat, LLMs show high pragmatic fluency (0.88-0.
arXiv

TAPe+ML v3 reaches 84.7 mAP50 and 65.3 mAP50-95 on COCO detection with under 100K parameters, and lifts 30-image industrial ore-blockage accuracy from about 28% with YOLO 26s to about 65%

The work presents TAPe+ML v3, which replaces raw pixel tensors with a structured TAPe (Theory of Active Perception) representation and combines background, pointer, and prototype-clustering submodels under a coordinator; with fewer than 100,000 parameters it reports 84.7 mAP50 and 65.3 mAP50-95 on COCO detection, 80.7 mask mAP50 and 58.4 mask mAP50-95 on COCO instance segmentation, 88.1% ImageNet-1k Top-1 and 89.9% ImageNet-Real, 92% versus 47% on Imagenette under identical training with TAPe versus raw pixels, and an industrial ore-blockage pilot where backbone adaptation outperforms head-only adaptation.
Nature News

Many Ways for a Cell to Die: From Apoptosis and Necroptosis to Alkaliptosis and Reversible Death

Drawing on a Nature news feature and a Nature Reviews Molecular Cell Biology review, this column-style summary surveys roughly 20 new cell-death modes described since 1999 — including alkaliptosis, pyroptosis, necroptosis, ferroptosis, cuproptosis, sodium-overload death and ruptosis — showing that how a cell dies shapes the signals it releases to neighbours, tissues and immune responses, and that death is not always irreversible, with 'bucket list' delays and anastasis revival.
arXiv

Realtime-Venus splits real-time dialogue from background work into a dual-loop runtime with two 9B frontends, leading online models on six of eight video benchmarks

The work presents Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models, Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction, which align continuous perception, conversational control, native speech generation, and private delegation requests on a shared causal timeline while Realtime-Venus-Harness executes background capabilities asynchronously in a dual-loop runtime, achieving the highest scores among compared online models on six of eight video benchmarks and leading or matching on several audio understanding, spoken question answering, and full-duplex overlap metrics.
arXiv

Aligning to Frozen World-Model Features Lifts a 0.8B Robot Policy to 97.9% on LIBERO and +2.3 Points on RoboCasa-GR1 at Zero Inference Cost

The work introduces a method that distills world-model representations into compact vision-language-action (VLA) policies: ordinary VLA training gains one cosine alignment term against cached features from a frozen world model, the teacher runs only once offline and is never loaded during training, the projector is discarded afterward, and the deployed policy is identical to the undistilled baseline; the 0.8B student reaches 97.9% average success across the four LIBERO suites (2.6 points above the identical undistilled student), improves RoboCasa-GR1 humanoid manipulation from 48.2% to 50.5%, and gains on both single-arm and bimanual real platforms, with the gain surviving changes of student scale, backbone, alignment layer, and teacher.
arXiv

SVEET trains only on a bidirectional video diffusion model and transfers editing zero-shot to a streaming backbone at 15 FPS on a single H100

The work proposes SVEET, a framework that trains a control branch only on a pretrained bidirectional video diffusion model and, through a temporally independent 2D spatial-attention control branch plus Orthogonal Decoupled Training (ODT) that bridges the bidirectional-to-autoregressive feature gap, enables high-quality streaming video editing without retraining or distilling the streaming backbone, reaching 15 FPS on a single H100 GPU and outperforming several baselines on style transfer, video inpainting, and depth-to-video generation.
arXiv

OmniEdu fine-tunes 4B/9B/27B models on 69,999 capability-oriented examples, beating each base model across curriculum grounding, K–12 problem solving, and pedagogical tutoring

Researchers from Peking University, the University of the Chinese Academy of Sciences, and Zhongguancun Academy built OmniEdu, an open family of K–12 learning-and-teaching foundation models, trained with a capability-oriented instruction corpus organized around subject competence, curriculum grounding, diagnostic reasoning, and pedagogical action and scaffolding (69,999 examples and 15.96M supervised response tokens, of which 60,951 examples and about 12.0M tokens are education-specific), fine-tuning 4B, 9B, and 27B backbones with full-parameter supervised fine-tuning and reporting consistent gains at every scale on curriculum-grounding, K–12 problem-solving, and pedagogical-tutoring benchmarks, with OmniEdu-27B reaching 63.12% EM / 76.69% F1 on K12-Bench, 85.89% on MathFish, 86.
arXiv

Jev-Mem puts a System-One control plane over agentic memory, lifting LoCoMo overall score to 0.777 while cutting memory build time to 158 seconds

Jev-Mem introduces an agentic memory architecture inspired by System-One/System-Two cognition, in which a dedicated System-One control plane handles memory typing, relation construction, query routing, retrieval-budget allocation, graph traversal, candidate scoring, and adaptive stopping, while System Two is reserved for complex reasoning and answer synthesis; on LoCoMo it reaches an overall LLM-as-a-Judge score of 0.777 (an 11.0% relative gain over the strongest baseline), a memory construction time of 158 seconds (a 6.6 speedup over the fastest competing memory system), and an average query latency of 0.93 seconds (a 36.7% reduction).
Nature News

Ultra-cold microscopy reaches the lab: the liquid-helium STEM GAIA and the road to atomic-resolution cryogenic imaging

This evidence bundle centres on a Nature news feature reporting that Bruker will ship GAIA, described as the first scanning transmission electron microscope to operate stably near absolute zero with liquid helium, to three laboratories in Canada, Germany and the United States, where it can image materials at about −266 °C (7 kelvin) for 30 hours or more, while tracing the earlier route: a 2025 US-reported TEM attachment reaching −253 °C (20 kelvin) with atomic-resolution imaging stable for more than ten hours, a condenZero liquid-helium attachment at around 4 kelvin for 24 hours that did not aim for atomic resolution, and Idrobo's 2021 comment on aberration correction combined with further instrument developments.
arXiv

Why Video Diffusion Models Break Physics: Researchers Trace It to RoPE Spatial Anchoring and Fix It by Rescaling RoPE Frequency

The work presents an interpretability study of the "motion planning" process in text-to-video diffusion models, finding that excessive spatial attention decay induced by RoPE in self-attention makes early candidate regions lock prematurely onto physically implausible positions, and proposes a lightweight architectural change that merely rescales RoPE frequency across denoising steps, improving physical commonsense in both training-free and training-based experiments.
arXiv

VideoGen-Agent uses multitask reinforcement learning to teach a video-generation agent to call retrieval, simulation, and verification tools, lifting its base generator from 56.5 to 75.6 on VABench and to 86.1 with upgraded generation tools

The work introduces VideoGen-Agent, a multimodal agent trained with supervised fine-tuning followed by multitask agentic reinforcement learning to coordinate augmentation, generation, and verification tools across six video-generation tasks (Procedural Knowledge, Single-Entity Identity, Multi-Entity Identity, Physics Simulation, Compositional Scene, and Multi-Shot), and builds VABench, a held-out benchmark of 600 prompts; on VABench it improves the base T2V generator Seedance 1.0 from 56.5 to 75.6, and further to 86.1 when compatible upgraded generation tools unseen during training are used, with human raters preferring that configuration in 100 pairwise comparisons.