Skip to main content

Daily report

AI and science frontiers · 2026-09-27

Only content delivered through the publication boundary on this date is included.

IEEE Spectrum

Retired IBM software engineer Ralph Earle turns code into interface beauty in the poem 'Poetry for Engineers'

This short piece presents 'Poetry for Engineers: The UI Designer's Dream,' a poem by Ralph Earle, a former technical editor and later senior software engineer at the IBM Software Lab, in which writing code and building the internet pixel by pixel is framed as translating dry code into surpassing beauty, and the speaker dreams of translating holy visions into a communication protocol of universal wonderment and letting shreds of light fall like diamonds on the endless mountains of the Web; an accompanying biography notes that Earle joined the IBM Software Lab in Raleigh, N.C., in 1990, retired as a senior software engineer in 2017, coauthored Enterprise Computing with Objects, and after retirement published two poetry collections and was nominated for the Pushcart Prize three times.
bioRxiv

Team builds a single-nucleus multi-omic atlas across 21 adult human tissues, profiling 459,856 transcriptomic and chromatin accessibility profiles and identifying 161,270 novel regulatory elements

The study presents a single-nucleus multi-omic atlas comprising 459,856 transcriptomic and chromatin accessibility profiles from 21 adult human tissues and four donors, including paired measurements from 160,688 nuclei, resolving nine cell lineages, 61 broad cell types and 313 subclusters, identifying 1,085,062 candidate cis-regulatory elements (including 161,270 novel elements absent from ENCODE), and using the dataset to train sequence-to-function models that predict chromatin-accessibility effects for 548,656 fine-mapped variants, identifying 18,133 high-effect variants including 1,120 broadly active variants.
发表出处待核验

A PhD thesis proposes three methodological innovations and a design framework for bringing causal machine learning into clinical practice, testing continuous treatment models on a physiotherapy dosage case

This PhD thesis addresses how causal machine learning can be translated into real-world clinical practice along three lines: for binary treatment effect estimation in clinical contexts it identifies eight key decision points that directly influence treatment effect estimates and proposes a structured framework for designing causal models in clinical settings; for the difficulty of verifying counterfactual predictions it proposes a novel method that grounds average model predictions in a group-level quantity that can be empirically verified using randomized or real-world clinical trial data, and shows that a model with a lower counterfactual error bound can be constructed and that information bottleneck regularization improves treatment effect estimation; for continuous causal machine learn
International Journal of Modern Science and Research Technology

Narrative review: GenAI cognitive offloading can boost immediate performance but shows a performance-learning dissociation, and which functions are offloaded matters more than whether AI is used

This narrative review searched peer-reviewed GenAI studies published through September 05, 2026 alongside foundational cognitive-offloading research and found that evidence does not support a simple positive or negative effect of GenAI on human cognition: GenAI-supported offloading of functions such as memory retrieval, processing, and learning can reduce required cognitive capacity and support high performance across a variety of tasks, yet the same offloading can in some cases reduce the quality of subsequent encoding, and recent GenAI studies show a consistent performance-learning dissociation in which gains from support available at the time typically carry over to later unsupported performance only slightly, not at all, or even negatively.
发表出处待核验

Doctoral thesis pooling homecare data from six countries links moderate-to-higher physical activity to better function, health outcomes and mortality in older homecare recipients

This doctoral thesis, conducted within the EU-funded I-CARE4OLD project, first maps the definition and use of non-pharmacological interventions (NPIs)—a systematic review shows highly heterogeneous and predominantly negatively framed definitions, and data from six countries show underuse of interventions with proven benefits (physical therapy, occupational therapy, psychosocial interventions, social participation) alongside relatively frequent use of physical restraints in some countries—and second, using longitudinal real-world data from the Netherlands and Ontario, Canada, with propensity-based methods, generalized estimating equations, and target trial emulation with causal machine learning, finds that moderate and higher levels of physical activity are associated with beneficial effect
IJIS - Indonesian Journal On Information System

Across seven algorithms for SPKLU review sentiment on PLN Mobile, MLP wins with 89.16% accuracy but 59.42% macro F1

The study collected user reviews of the EV Charging/SPKLU feature in PLN Mobile from Google Play Store, labeled sentiment using ratings as weak labels, applied algorithm-specific class weighting on the training data to address class imbalance, and compared seven classifiers—Logistic Regression, SVM, Naive Bayes, MLP, RNN, BiLSTM, and DistilBERT—finding that MLP achieved the highest validation macro F1 (67.80%) and was selected for final testing, where it reached 89.16% accuracy and 59.42% macro F1 on 286 test instances, indicating high overall classification performance with uneven performance across sentiment classes.
JURNAL PESISIR DAN LAUT TROPIS

Machine-learning bias correction cuts Ina-Flows SST forecast error by 29.83% at Karimun Jawa and 26.10% at Bira, with the best algorithm differing by site

Using MAWS observations, this study evaluated the bias in BMKG Ina-Flows sea surface temperature forecasts and compared three machine-learning bias-correction algorithms (SVR, LSTM, Bi-LSTM) at Karimun Jawa and Bira, finding a warm bias at both sites, with LSTM best at Karimun Jawa (RMSE 0.207, MAE 0.182, MBE -0.171, a 29.83% RMSE reduction) and Bi-LSTM best at Bira (RMSE 0.218, MAE 0.173, MBE 0.011, a 26.10% RMSE reduction), indicating that correction effectiveness is location-dependent and algorithm choice should rest on local validation.
arXiv

Reflex-Guard cuts prompt-safety filtering to 37.6 ms locally with dense semantic embeddings, reaching 95.9% harmful-prompt recall on 30,568 samples

The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
arXiv

TP-CRIV lets a third party with no white-box or API access verify an AI model's identity using only its ordinary black-box inference interface

The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a setting where the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under fresh, previously undisclosed requirements and network isolation it obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models showing clear same/cross-model separation and finite-chal
arXiv

LiveMathematicianBench builds a research-level math reasoning benchmark from fresh arXiv papers, where the best model reaches only 43.5% and Gemini-3.1-pro-preview falls to 17.6% under substitution-resistant evaluation

The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs, introducing a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism; evaluation shows the benchmark is far from saturated, with the best model Gemini-3.1-pro-preview reaching only 43.5%, substitution-resistant evaluation yielding a top score of 30.6% for GPT-5.4 while Gemini-3.1-pro-preview drops to 17.6% below the 20% random baseline, and a dual-mode protocol showing consistent accuracy gains from proof-sketch access.
arXiv

DMA approximates factor-to-variable messages directly, bounds marginal KL by message KL, and yields a Bayesian neural network inference algorithm without a learning-rate hyperparameter

The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
arXiv

SciMuse generates personalized research ideas from a 58-million-paper knowledge graph and an LLM; over 100 research group leaders rated 4,400+ ideas at a mean of 2.40 on a 5-point scale, with 24.9% rated 4 or 5

The authors introduce SciMuse, which generates personalized research ideas using a knowledge graph of 58 million papers and a large language model, and they conducted a large-scale evaluation in which more than 100 research group leaders spanning the natural sciences to the humanities rated over 4,400 personalized ideas by level of interest; expert ratings were modest (mean 2.40 on a 5-point scale, most common rating 1) while 24.9% of ideas were rated 4 or 5, supplying knowledge-graph-selected concept pairs did not improve expert-rated interest over a titles-only GPT baseline, high-citation-predicted pairs even showed a weak tendency (1.
arXiv

MorphIK solves inverse kinematics for unseen 6-to-9-DoF robots with morphology-conditioned flow matching, reaching about 5 cm and sub-millimeter accuracy after optimization

MorphIK is a flow-matching model that encodes a robot's morphology together with the target pose using a transformer and conditions a flow-matching head on that encoding to generate joint poses from noise, thereby solving inverse kinematics for revolute-joint kinematic chains never seen during training; trained only on purely synthetic data from procedurally generated robots, it reaches about 5 cm precision on unseen real-world robots with 6 to 9 Degrees of Freedom, serves as a prior that reduces error to less than 1 cm after a single step of Damped Least Squares optimization and to sub-1 mm after 3 steps in most cases, and can efficiently sample the null space to yield varied configurations for the same pose.
arXiv

NOVQS combines multiple shallow parameterized quantum circuits and matches or outperforms a single deeper circuit in hydrogen-chain and nitrogen-molecule simulations

The work introduces nonorthogonal variational quantum simulation (NOVQS), which applies linear combinations of parameterized quantum states to real- and imaginary-time evolutions, designs a shallow hardware-friendly ansatz tailored to second-quantized electronic-structure Hamiltonians together with resource-efficient protocols for measuring the matrices and vectors in the parameter equations of motion, and provides error analysis and resource estimation; numerical simulations of hydrogen chains and the nitrogen molecule show that a collection of shallow, or even single-layer, parameterized quantum circuits can match or outperform a much deeper circuit in variational quantum simulation, revealing a trade-off between circuit number and depth.
arXiv

ProGS organizes formal modeling around model-based proof sketches, improving syntactic validity, deductive verifiability, and behavioral correctness on a 27-system benchmark

The work proposes Proof-Sketch-Guided Formal Model Synthesis (ProGS), an autoformalization method centered on model-based proof sketches: a sketch represents the proof structure of the target formal system as a tree, with internal nodes capturing case splits and inductive reasoning steps and leaf nodes corresponding to concrete state-transition events that realize individual subgoals; LLMs generate and repair these sketches, and verification failures are mapped back to specific nodes and subtrees to provide structured guidance for iterative repair; evaluation on a benchmark of 27 formal systems shows ProGS improves over state-of-the-art agentic formal modeling approaches in syntactic validity, deductive verifiability, and behavioral correctness.
arXiv

A new CDP diagnostic finds DeepPersona-Inspired prompting shows the strongest cultural flattening in every model-domain block across four backbones and two survey domains

The work introduces Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration that flags reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature, and evaluates it across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test; results reveal a systematic discrepancy between conventional fidelity metrics and CDP, with controlled experiments showing CDP changes monotonically as cross-country divergence is attenuated or amplified while corresponding JSD changes remain relatively small, and an audit of real LLM generations showing DeepPersona-Inspired prompting is frequently favored by conventional fidelity m
arXiv

ArGuard Shared Task: 58 teams registered and 35 competed on Arabic meme and LLM-prompt harm detection, with best systems reaching macro-F1 of 0.823/0.419/0.984/0.790

ArGuard is a shared task on harmful content detection in Arabic, with two tracks: Track A on multimodal hate detection in Arabic memes and Track B on harmful prompt detection for Arabic LLM safety evaluation; 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers, with teams exploring models such as AraBERT, Jais, and Qwen3-VL, and the best systems achieving macro-F1 of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2, where fine-grained meme classification in A2 was the most challenging setting due to sparse labels and train-test distribution shifts.
arXiv

Review maps how large language models enter nanophotonic design as surrogate models and agentic systems

This review surveys how large language models add semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows, organizing the methods into two operational modes: surrogate models that treat structure-spectrum mapping as a language task, and agentic systems demonstrated to generate code, orchestrate selected simulation steps, and support closed-loop optimization; it also traces the development from classical neural networks to transformer-based models and briefly explores cross-disciplinary applications in fields such as materials science and wireless communications, looking ahead to next-generation multimodal foundation models with physical perception.
arXiv

Twenty-seven middle school teachers configured pedagogical intent in a teacher-facing chatbot authoring tool, and log-based evaluation showed 88.9% alignment for responsiveness and 81.5% for persona versus 70.4% for rules and 59.3% for purpose

In professional development workshops, 27 middle school teachers used a teacher-facing chatbot authoring tool, and by analyzing focus-group interviews alongside configuration and interaction logs the study examined how pedagogical intentions are translated into chatbot configurations and reflected in chatbot behavior, finding that teachers envisioned chatbots as instructional scaffolds offering differentiated support, extending access to assistance, and preserving student thinking within teacher-defined boundaries; configuration analysis showed Purpose primarily captured instructional goals and content focus while Rules more often specified pedagogical behavior, guardrails, and learner-specific adaptations; and log-based evaluation showed stronger alignment for responsiveness (88.
arXiv

Learning-based framework enables continuous autonomous excavation on a scaled hydraulic excavator, averaging 6.52 kg payload per cycle versus 2.68 kg for Fixed Dig

The work presents a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers: a shared task-conditioned RL policy handles waypoint-guided approach and loaded transport, an IL policy learns vision-based digging and lifting from expert demonstrations, and digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints; deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control, offline replay and physical experiments show more consistent target selection, shorter local motion time, and increased payload, with the learned digging policy achieving a mean payload of 6.52 kg per completed cycle versus 2.
arXiv

TabSieve's select-then-predict framework lifts classification by 2.92% and regression by 4.45% on average across 75 classification and 52 regression tables

The authors propose TabSieve, a select-then-predict framework in which, given a table and a query row, the model first selects a small set of informative rows as evidence and then predicts the missing target conditioned on that evidence; to enable this capability they build TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables with a strong teacher model and strict filtering, and introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards and stabilizes mixed regression and classification training via dynamic task-advantage balancing; on a held-out benchmark of 75 classification and 52 regression tables, TabSieve consistently improves performance across shot budgets, with
arXiv

CANOPY uses random-path probes to certify smoothness violations online, improving routing, top-k identification, test-time search, caching, and prompt trimming at matched budgets

The work introduces CANOPY, a multi-fidelity tree bandit that uses cheap random-path probes to build an online certificate of local aggregation bias and directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation, rather than assuming a global smoothness prior; the authors prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense; across routing, top-k identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including 2.9x higher top-10 recall on a 1000-model pool, 1.
arXiv

GPT models change more lines than human patches when fixing Codeforces submissions, and solve more problems from scratch than by patching

This work collects roughly 3000 submissions from a couple of Codeforces users, pairs each buggy submission with its corresponding human fix, uses the similarity between the buggy solution and the human fix as a baseline, evaluates the quality of LLM-generated bug fixes on three OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1), and checks whether generated solutions solve the problem using the Codeforces-R1 dataset; the findings suggest LLMs tend to modify more lines than human fixes and sometimes generate entirely new solutions, and that LLMs solve more problems correctly when allowed to generate from scratch rather than patch buggy submissions, even when those submissions are close to the human patch.
arXiv

STAM-ASR extends a pretrained AudioLLM with speaker-temporal anchoring and memory, evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 for multi-speaker ASR

The work proposes STAM-ASR, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR without relying on an external diarization system, learning speaker activity and speaker-aware representations directly from intermediate AudioLLM features to provide explicit who-and-when cues that modulate the semantic representation, while maintaining fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits and that the gap between reference and predicted speaker activity identifies robust speaker tracking
arXiv

Does per-frame early exit pay off? A dynamic-depth speech enhancer lands on the same latency-quality frontier as static models on STM32N6, with a 26 microsecond per-frame policy

The work supervises every intermediate depth of one causal speech enhancement model and fine-tunes its output heads so deeper outputs are never worse than shallower ones, yielding a family of static models that are more Pareto-efficient than equivalently-sized counterparts trained from scratch on the same budget (up to 0.11 higher PESQ at equivalent compute, matching the best PESQ at 30% less compute); after int8 quantization on an STM32N6 microcontroller, the dynamic enhancer lies on the same latency-quality frontier as the static models, with the policy running in 26 microseconds per frame on the companion Cortex-M55 and splitting the enhancer into separate NPU graphs adding 2.2% latency overhead.
arXiv

A frozen Sapiens pose foundation model detects climbing hold usage without climbing-specific training, reaching 90.2% event F1

The work proposes a training-free method that uses fingertip and toe keypoints from a frozen Sapiens pose foundation model, combined with a per-frame proximity test against annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, to detect which holds a climber uses and when; on the The Way Up dataset (22 videos, 10 athletes, two routes) it reaches an event F1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, performs best on footholds (F1 89.8% overall, 96.
arXiv

An empirical analysis of 2,042 Python repositories finds 11.2% show OS-dependent test failures, yielding a 7-category taxonomy and an LLM repair evaluation

This study presents a large-scale empirical analysis of cross-OS portability issues across 2,042 open-source Python repositories, combining cross-OS test reexecution (500 projects, 11.2% showing OS-dependent test failures) with manual analysis of 240 GitHub issues (confirming 102 genuine portability problems across 95 additional projects), and develops a taxonomy of 7 primary categories, 24 sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns, while evaluating existing static analysis tools as providing minimal support and large language models achieving 40-79% accuracy in identifying issues and 50-77% success in generating fixes under structured guidance, with practical applicability demonstrated through 33 contributed pull requests (17 merged, zero rejected).
arXiv

After deploying a persistent, proactive AI 'teammate' across multiple teams at a large technology company, researchers observed human-agent workplace boundaries being continually negotiated across collaboration rules, the non-human actor's relational boundaries, and the redistribution of trust.

This paper presents an in-situ qualitative study of a persistent, proactive AI agent 'teammate' deployed across multiple teams in a large technology company, finding that the boundaries of the human-agent workplace are actively in flux and trigger breakdowns and negotiations across three areas: tacit rules of collaborative human workflows, the relational boundaries of this new non-human actor, and the redistribution of trust and human agency; the authors use these early micro-negotiations as signals to chart a research, design, and organizational agenda that intentionally preserves human agency.
arXiv

Agent authorization recast as a cryptographically verifiable relation R_CVA, with an executable Groth16 zk-SNARK proof of concept

The study hypothesizes that agent authorization can be formalized as a cryptographically verifiable relation, R_CVA, that jointly binds an agent principal, a concrete authorization request, an execution context, and the satisfaction of an applicable policy while selectively preserving the confidentiality of private authorization attributes; it introduces a preliminary formal abstraction for Cryptographically Verifiable Agent Authorization (CVA), defines a compact set of candidate security properties including authorization soundness, principal binding, request binding, policy binding, and replay resistance, provides an executable zero-knowledge proof of concept instantiating selected elements of the model over a Groth16 zk-SNARK construction, and formalizes the structural separation among
arXiv

MemGuard-Alpha audit finds MIA contamination scores drop to 0.487–0.537 discrimination at a fixed date, and no filtering variant beats the unfiltered ensemble under symmetric transaction costs

The authors introduce MemGuard-Alpha, comprising a composite contamination score combining five membership inference attack (MIA) methods with a temporal proximity feature, plus Cross-Model Memorization Disagreement exploiting variation in training cutoffs across models, and audit both across seven LLMs (124M–7B), 50 S&P 100 constituents, 42,800 prompts and 299,600 prompt-model MIA scores spanning 2019–2024, yielding three negative results: the temporal proximity feature recovers the in-sample label perfectly (ROC-AUC 1.000) because it is a monotone transform of the defining variable, the discriminative power of MIA scores is largely attributable to scale differences between models (raw scores reach AUC up to 0.99 at a fixed date but fall to 0.487–0.
arXiv

A 146-junction single-walled carbon nanotube library shows mean chiral angle sets the averaged transmission step while metallic or semiconducting character sets the junction gap

The work builds a computational library of 146 single-walled carbon nanotube (SWCNT)–SWCNT junctions and analyses their magnetotransport with an automated workflow combining molecular dynamics, tight-binding theory, Peierls magnetic coupling, and non-equilibrium Green's functions, followed by machine-learning analysis; it finds that the averaged first transmission-step value is governed primarily by the mean chiral angle of the two nanotubes, that the junction energy gap depends predominantly on the metallic or semiconducting character of the constituent nanotubes, that temperature generally suppresses the averaged transmission while reducing the extracted gap, that a perpendicular magnetic field affects transmission much more strongly than the gap, that signatures of interference-driven t
arXiv

Reframing audio description as constrained optimization over what, when, and how lets a hybrid LLM-plus-MILP system set new state of the art on REFRAMED's narrative QA and temporal metrics

The work formalizes audio description (AD) generation as a constrained optimization problem over three coupled decisions—what visual information to describe, when it can be spoken in gaps between dialogue, and how to formulate it to fit the available time—and proposes a hybrid system in which large language models propose and ground visual elements, estimate their narrative salience, and generate compressed realizations, while a mixed-integer linear program jointly selects and schedules descriptions across a scene under temporal constraints; evaluated on REFRAMED, a benchmark for realistic AD of movies, the approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new state of the art on narrative QA and temporally grounded metrics, w
arXiv

OncoVision's attention-driven multimodal training framework cut reading time by up to 61% and raised diagnostic confidence in a paired six-radiologist evaluation

OncoVision is a privileged-information training framework that uses mammography images and clinical features during training while performing inference from mammographic images alone; built on an attention-based encoder-decoder backbone, it jointly segments four regions of interest (masses, calcifications, axillary findings, and breast tissue) with accuracy exceeding the nnU-Net baseline and predicts ten structured clinical features including BI-RADS category; the authors developed two late-fusion strategies, Independent and Dependent, that integrate imaging, radiomic, and clinical information during training, with radiomic features extracted from predicted masks providing shape, intensity, and texture descriptors that complement the learned CNN representations; in a retrospective multi-re
arXiv

SkillGym turns human skills into verifiable training environments, letting a 35B model reach 51.47% on SkillsBench and exceed reported scores of Claude Sonnet 4.6 and others

SkillGym introduces a framework that transforms human-written agent skills into executable, verifiable training environments: its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions; it constructs and releases 2,756 environments across 12 categories and collects 8,364 successful trajectories (averaging 49 tool calls and over 60k logged text tokens), supporting supervised fine-tuning on verified workflows and reinforcement learning with outcome-based rewards; under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.
arXiv

STAM-ASR extends a pretrained AudioLLM with speaker-temporal anchoring and memory for multi-speaker ASR, evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1

STAM-ASR is a lightweight framework that, without relying on an external diarization system and without explicit speech separation, learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features to give explicit who-and-when cues that modulate the AudioLLM's semantic representation, and maintains fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
arXiv

TIDE reuses inference-time hidden states to adapt draft models online, reaching up to 1.66x throughput and recovering performance when static drafts degrade it

TIDE is a serving-engine-native framework that reuses the target model's intermediate hidden states produced during inference as training signals for online draft adaptation, avoiding additional target model computation and serving-time overhead, and combines this with adaptive runtime control that activates speculation and draft training only when beneficial and with mapping of inference and training onto appropriate GPU classes; across diverse real-world workloads TIDE achieves up to 1.66x throughput over no-speculation baselines, recovers performance on misaligned workloads where static draft models degrade throughput, reduces training time by up to 3.02x and storage requirements by 24x compared to existing draft training approaches, and improves system throughput by up to 1.
arXiv

SEEK externalizes search evaluation criteria into a routable skill bank, improving listwise quality evaluation and attribution diagnosis in Kuaishou short-video search

The work proposes SEEK (Skill-routed Evaluation with Evolvable Knowledge), which externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution; a two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining; experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis, and SEEK has been deployed at Kuaishou, a platform with over 400 million daily activ
arXiv

Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model in tool use, memory, skills, and sub-agent coordination

The work builds Qwen-Planner-Agent within a closed-loop AI-for-AI framework that links data production, model training, and deployment through a shared action-feedback-verification contract: AI for Data builds a human-gated agentic data flywheel, AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning and introduces CARE to reduce reasoning and tool-use costs, and AI drives model-harness co-evolution; the agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination, with further gains on non-mobile agentic benchmarks while largely preserving general capabilities.
arXiv

Unite-Audio jointly trains audio representation learning with latent flow matching, reaching competitive text-to-audio generation with a compact model

The work introduces Unite-Audio, which jointly learns continuous audio representations and latent flow matching for text-to-audio generation; by coupling reconstruction with self-supervised generative prediction it lets the generative objective directly shape the latent space, and it applies Flow-GRPO post-training to improve text-conditioned generation; experiments show competitive text-to-audio performance with a compact latent flow model, and ablation studies confirm the benefit of jointly learning the audio representation and the generative model.
arXiv

Zero-Data Self-Play Pretraining: a generator and a learner trained in tandem from random initialization show zero-shot loss scaling predictably with self-play compute across several natural datasets

The work introduces Self-Play Pretraining with Zero Data as an initial proof-of-concept: starting from random initialization, a generator proposes programs interpreted by a universal Turing machine to produce byte sequences while a learner autoregressively predicts those byte sequences with standard cross-entropy, and the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum; because neither generator nor learner is trained on natural data, zero-shot performance on natural data serves as a clean test of transfer, and across several natural datasets zero-shot loss exhibits predictable scaling in compute, with the models also exhibiting in-context learning and discovering recognizable mathematical
arXiv

MILO cuts many-shot context KV cache by up to 50% via block-wise low-rank compression, lifting Qwen2.5 throughput 1.8x with near-flat classification and reasoning performance

The work proposes MILO, a compression framework that exploits low-rank redundancy in many-shot contexts by compressing the KV cache at block granularity and dynamically allocating rank budgets based on information entropy, achieving up to 50% KV cache memory reduction and 1.8x throughput improvement on Qwen2.5 models with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.
arXiv

Full development history of a wholly AI-authored codebase released: 14.3% of AI code-generation events in a 21,000-line Python tool contained real errors, and roughly 1 in 4-5 interactive responses contained factual errors

The work releases a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI with no human-authored code or tests, together with two code-provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability; applying these to the dataset, it finds that user coding agent CLI instructions differ in kind from IDE-chat instructions with a greater focus on comprehension, planning and consultation, that code development is mainly proactive, that 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, and that roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.
arXiv

WebArxiv evaluates web agents on 510 static-snapshot tasks, finds reliance on fixed interaction histories, and adds a dynamic-memory mechanism

The work introduces WebArxiv, a static-snapshot benchmark built on arXiv with 510 time-invariant tasks, each with a unique deterministic ground truth, for evaluating multimodal web agents on scholarly tasks such as multi-constraint paper retrieval, fine-grained content extraction, and cross-paper comparison; evaluations of a range of foundation-model-based web agents show the benchmark remains challenging, behavioral analysis reveals that agents over-rely on fixed interaction histories, causing incomplete or repetitive reasoning, and the authors therefore equip agents with a lightweight dynamic-memory mechanism for adaptive retrieval and reasoning over relevant context.
arXiv

Reading answer-label logits under a chain-of-thought cue drops Qwen2.5-VL-7B on ScienceQA from 80.76% to 45.48%, with 93.54% of predictions landing in the first option slot

The work identifies an evaluation practice it calls CoT-prefix scoring, in which a reasoning cue is appended to the prompt but answer-label logits are read before the model generates any rationale, and reports that on ScienceQA Qwen2.5-VL-7B falls from 80.76% to 45.48%, that 93.54% of CoT-prefix predictions select the first slot across five option-content permutations, and that condition-matched linear probes recover 78.94% from the same hidden states while free generation restores 75.24%, with vocabulary and layer diagnostics showing probability mass shifting toward continuation tokens while answer information remains linearly accessible in late layers, indicating an evaluation-interface mismatch rather than missing model knowledge.
arXiv

Audit finds a multilingual affective generation benchmark's headline conclusions stem from the measurement instrument, not system differences, and proposes an emoji-affect decodability probe

This work audits a multilingual affective generation benchmark in which eight instruction-tuned LLMs produced emoji summaries for 17,100 Bangla, English and Hindi sentences with 6,960 human judgements, finding that when annotators are treated as a random rather than a fixed factor no system differs significantly from any other (F(7,14)=0.59, p=0.76) although the conventional analysis declares 19 of 28 pairwise differences significant; annotator identity explains far more rating variance than system identity and removing any single annotator changes the winning system; the ordering that emerges tracks output length, with mean emoji count explaining 78.
arXiv

Handing governance review gates to agents: top model reaches 94.98% strict gate success on DGF-Bench but only 76.92% complete-route success

The work develops a task-substitution framework for Digital Governance Frameworks (DGF) that treats each governance gate as an executable contract, requiring sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted; it derives a residual-work threshold showing why automating most cases can still increase labor, and measures models on DGF-Bench (300 synthetic projects, 899 evaluable model-project runs): Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%, with complete-route success of 76.92%, 42.33%, and 24.
arXiv

Pain location's diagnostic value is split into three failures—anatomical multiplexing, central amplification, and person-dependent displacement—not one gradient

The paper argues that patient-reported pain location is sometimes diagnostically decisive and sometimes nearly uninformative because it contains three epistemically distinct failures—anatomical multiplexing as a non-identifiable inverse problem, delocalized amplification (clinically central sensitization or nociplastic pain) as a change of generative model, and referred and atypical displacement hypothesized as a systematic, person-dependent shift—which are one Bayesian inference problem failing at the likelihood, the model class, and the group-conditional prior, with a fourth node at the report itself; on this basis the author finds that the published "high-utility" accuracy band leans on overstated specificity, so the gradient is real but flatter than drawn, and attributes the finding th
arXiv

Reflex-Guard achieves 95.9% harmful-prompt recall at 37.6 ms locally, faster than Llama Guard 2 and SafeDecoding

The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
arXiv

XACT learns sparse attribution masks over invertible time-frequency transform coefficients, highlighting fewer spurious features than baselines on synthetic data and yielding sparse, structured explanations on two real-world datasets

The authors propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms (STFT, continuous wavelet transform, discrete wavelet transform), and extend the virtual inspection layer from the STFT to both wavelet transforms so that LRP can generate explanations in these representations; on a synthetic dataset XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines, and across two real-world datasets it produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria.
arXiv

TP-CRIV proposes a third-party challenge-response identity verification framework, achieving separable same-model and cross-model verification on ten ImageNet-pretrained TorchVision models

The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a third-party setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under these constraints the framework obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models demonstrating clear same/cross-model separation
arXiv

LiveMathematicianBench evaluates LLMs on arXiv theorems published after training cutoffs: best model reaches only 43.5%, and Gemini-3.1-pro-preview falls to 17.6% under substitution resistance, below the random baseline

The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs; it introduces a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism, and evaluation shows the benchmark is far from saturated—the best model, Gemini-3.1-pro-preview, achieves only 43.5%, while under substitution-resistant evaluation GPT-5.4 scores highest at 30.6% and Gemini-3.1-pro-preview falls to 17.6%, below the 20% random baseline; a dual-mode protocol shows proof-sketch access yields consistent accuracy gains.