Skip to main content

Search

“All disciplines” · 1494 results

Page 33 · showing 20
arXiv

SAGE combines algebraic sparsification and hyperbolic structural guidance to curb long-horizon reasoning biases, beating baselines across 12 benchmarks and 7 model families with up to 8-fold gains on Andrews-Curtis

The work introduces Symbolic Closure Analysis (SCA) to characterize exploration bias and compounding bias in long-horizon reasoning, and builds SAGE, a framework that injects structural priors during post-training via algebraic sparsification and hyperbolic structural guidance, outperforming SFT, GRPO, EMPO, and GRPO-PRM across 12 benchmarks and 7 model families and achieving up to an 8-fold improvement in Lean-verified proofs on the open Andrews-Curtis task.
arXiv

ZooWork aligns e-commerce rerankers with cross-family LLM judge labels: the 8B and 4B significantly beat the strongest open baseline on ShopRank-Bench, and the distilled 0.6B matches a 4B base on structured text

ZooWork presents ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, 8B) trained on preference labels produced by three reasoning LLM families (Qwen3.5-122B, Gemma-4-31B, DeepSeek-V4-Pro) under a constraint-first, both-orders judging protocol, with the aligned 8B distilling into the 4B and 0.6B; on ShopRank-Bench, built from Gensmo private traffic (10,511 pairs tiered gold/silver/bronze), the 8B and 4B significantly outperform the strongest open baseline Jina-m0, every size significantly beats its own un-aligned base, the 0.6B is statistically indistinguishable from the 4B base on structured text, and alignment costs nothing in general MTEB reranking quality or serving latency.
arXiv

InternW0-Δ pretrains a world action model on 20K+ hours of heterogeneous data, reaching 92.8% on LIBERO-Plus and 71.9% on RoboTwin 2.0 Clean2Random

InternW0-Δ couples a pretrained video expert and an action expert through a directed Mixture-of-Transformers, uses a frozen VLM for scene semantics, learns future-relevant scene changes via Causal Imprint from training-only future supervision, and distills 4D geometric and motion priors from a Track4World teacher at training time only; pretrained on a corpus of over 20K hours unifying robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data under a canonical state-action representation, it reaches 92.8% on LIBERO-Plus, 71.9% on RoboTwin 2.0 Clean2Random, an overall score of 66.0 on EBench, and an average success rate of 23.91% on RoboDojo, with deployment on four real-robot platforms (two gripper-based, two dexterous-hand).
arXiv

Across 2,170 GitHub projects, Jev drew 1,865 new repos in one week, yet 63% of attention went to routing and interface agents

The study presents a large-scale, data-driven analysis of 2,170 publicly available Jev projects collected from GitHub as of September 22, 2026, finding rapid early growth with 1,865 new repositories and 305 integrations into existing repositories within a week of release, projects using Jev for multiple decision purposes and combining its interfaces across domains, with attribute judgment (77%) and scoring or ranking (52%) most common, while public attention concentrates in routing and interface agents (19.6% of projects but 63.0% of stars) and does not track project counts.
arXiv

ARGUS audits identification assumptions in climate-policy DID studies, detecting 8 of 11 injected flaws and abstaining on about 60% of paper-dimension assessments across 26 economics papers

The authors introduce ARGUS, a bounded retrieval-gated large-language-model pipeline that decomposes a difference-in-differences (DID) study into eleven assumption-implication-evidence dimensions, audits whether the evidence a paper reports adequately supports each identification assumption, and abstains when no relevant evidence can be retrieved; it detects 8 of 11 planted flaws (versus 2 for a keyword baseline), 25 of 33 flaw variants, abstains on roughly 60% of paper-dimension assessments over 26 economics papers for lack of retrievable evidence, and in a five-paper, 55-cell pilot with two reconciled annotators is more severe than the human labels on 25 of the 33 cells it completes, an over-severity a rule fixed before the labels arrived reduces substantially in-sample.
arXiv

CDB lets looped language models infer at token-adaptive depth, realizing about 99% of the early-exit speedup available on Ouro 1.4B

The work introduces and first implements end-to-end continuous depth batching (CDB) for looped language models: it splits inference into prefill, prelude, recurrent-core, and coda stage queues, reforms batches between loop steps, manages a depth-indexed KV cache, and uses a lookahead gate to predict exits one step ahead so tokens with different loop depths can share a forward pass; on Ouro 1.4B and Huginn 3.5B, CDB realizes about 99% and 83%–96% of the estimated maximum speedup, showing fully looped architectures are best suited to depth-adaptive inference.
arXiv

SLCA-GRPO routes tool-segment and summary-segment advantages separately, raising tool-calling success by up to 10.03 points across three backbones

The work splits tool-calling RL trajectories into a tool segment and a summary segment, argues that broadcasting a single trajectory-level advantage causes cross-segment credit misattribution, and proposes SLCA-GRPO, which normalizes the two segment rewards within the group and routes each only to its own tokens inside a single unified policy, paired with a Schema-Guided LLM Simulator and Hierarchical Rewards; across Qwen2.5-3B/7B-Instruct and Qwen3-8B-Base on Toucan-Test, BFCL V3 and τ²-Bench it reports higher means than matched GRPO, with 7B gaps of +2.53, +1.36 and +9.15 percentage points.
arXiv

FoMo turns the forking moment of a diffusion trajectory into automatic perceptual-distance labels, reaching 0.733 SROCC on PIPAL and beating metrics trained on human-annotated data

The work proposes FoMo: forward-noising a reference image to a sampled timestep and then denoising it independently, using that forking moment as a pointwise perceptual-distance label, so training data can be generated with no human annotation, and trains reference-based IQA metrics with a RankNet-style global ranking objective; across four benchmarks and seven backbones it achieves the best average performance, with LPIPS-Alex reaching 0.733 SROCC on PIPAL versus 0.577 for the same architecture trained on human-annotated KADID-10k labels.
arXiv

Tactile-JEPA pretrains on the taxel connectivity graph and cuts force-estimation error by 6.3% and in-hand pose error by 20.8%

The work introduces Tactile-JEPA, a self-supervised pretraining method for distributed tactile sensors (e-skins) that samples masks at local and global scales over the taxel connectivity graph and predicts the embeddings of masked taxels; across magnetic and piezoresistive sensors and three datasets (Sparsh-skin, Tactile socks, DECO-50) it reduces force-estimation RMSE by 6.3% and in-hand pose RMSE by 20.8% over the strongest prior baseline, with consistent gains in action classification, object classification, and tactile-conditioned policy learning.
IEEE Spectrum

Retired IBM software engineer Ralph Earle turns code into interface beauty in the poem 'Poetry for Engineers'

This short piece presents 'Poetry for Engineers: The UI Designer's Dream,' a poem by Ralph Earle, a former technical editor and later senior software engineer at the IBM Software Lab, in which writing code and building the internet pixel by pixel is framed as translating dry code into surpassing beauty, and the speaker dreams of translating holy visions into a communication protocol of universal wonderment and letting shreds of light fall like diamonds on the endless mountains of the Web; an accompanying biography notes that Earle joined the IBM Software Lab in Raleigh, N.C., in 1990, retired as a senior software engineer in 2017, coauthored Enterprise Computing with Objects, and after retirement published two poetry collections and was nominated for the Pushcart Prize three times.
bioRxiv

Team builds a single-nucleus multi-omic atlas across 21 adult human tissues, profiling 459,856 transcriptomic and chromatin accessibility profiles and identifying 161,270 novel regulatory elements

The study presents a single-nucleus multi-omic atlas comprising 459,856 transcriptomic and chromatin accessibility profiles from 21 adult human tissues and four donors, including paired measurements from 160,688 nuclei, resolving nine cell lineages, 61 broad cell types and 313 subclusters, identifying 1,085,062 candidate cis-regulatory elements (including 161,270 novel elements absent from ENCODE), and using the dataset to train sequence-to-function models that predict chromatin-accessibility effects for 548,656 fine-mapped variants, identifying 18,133 high-effect variants including 1,120 broadly active variants.
发表出处待核验

A PhD thesis proposes three methodological innovations and a design framework for bringing causal machine learning into clinical practice, testing continuous treatment models on a physiotherapy dosage case

This PhD thesis addresses how causal machine learning can be translated into real-world clinical practice along three lines: for binary treatment effect estimation in clinical contexts it identifies eight key decision points that directly influence treatment effect estimates and proposes a structured framework for designing causal models in clinical settings; for the difficulty of verifying counterfactual predictions it proposes a novel method that grounds average model predictions in a group-level quantity that can be empirically verified using randomized or real-world clinical trial data, and shows that a model with a lower counterfactual error bound can be constructed and that information bottleneck regularization improves treatment effect estimation; for continuous causal machine learn
International Journal of Modern Science and Research Technology

Narrative review: GenAI cognitive offloading can boost immediate performance but shows a performance-learning dissociation, and which functions are offloaded matters more than whether AI is used

This narrative review searched peer-reviewed GenAI studies published through September 05, 2026 alongside foundational cognitive-offloading research and found that evidence does not support a simple positive or negative effect of GenAI on human cognition: GenAI-supported offloading of functions such as memory retrieval, processing, and learning can reduce required cognitive capacity and support high performance across a variety of tasks, yet the same offloading can in some cases reduce the quality of subsequent encoding, and recent GenAI studies show a consistent performance-learning dissociation in which gains from support available at the time typically carry over to later unsupported performance only slightly, not at all, or even negatively.
发表出处待核验

Doctoral thesis pooling homecare data from six countries links moderate-to-higher physical activity to better function, health outcomes and mortality in older homecare recipients

This doctoral thesis, conducted within the EU-funded I-CARE4OLD project, first maps the definition and use of non-pharmacological interventions (NPIs)—a systematic review shows highly heterogeneous and predominantly negatively framed definitions, and data from six countries show underuse of interventions with proven benefits (physical therapy, occupational therapy, psychosocial interventions, social participation) alongside relatively frequent use of physical restraints in some countries—and second, using longitudinal real-world data from the Netherlands and Ontario, Canada, with propensity-based methods, generalized estimating equations, and target trial emulation with causal machine learning, finds that moderate and higher levels of physical activity are associated with beneficial effect
IJIS - Indonesian Journal On Information System

Across seven algorithms for SPKLU review sentiment on PLN Mobile, MLP wins with 89.16% accuracy but 59.42% macro F1

The study collected user reviews of the EV Charging/SPKLU feature in PLN Mobile from Google Play Store, labeled sentiment using ratings as weak labels, applied algorithm-specific class weighting on the training data to address class imbalance, and compared seven classifiers—Logistic Regression, SVM, Naive Bayes, MLP, RNN, BiLSTM, and DistilBERT—finding that MLP achieved the highest validation macro F1 (67.80%) and was selected for final testing, where it reached 89.16% accuracy and 59.42% macro F1 on 286 test instances, indicating high overall classification performance with uneven performance across sentiment classes.
JURNAL PESISIR DAN LAUT TROPIS

Machine-learning bias correction cuts Ina-Flows SST forecast error by 29.83% at Karimun Jawa and 26.10% at Bira, with the best algorithm differing by site

Using MAWS observations, this study evaluated the bias in BMKG Ina-Flows sea surface temperature forecasts and compared three machine-learning bias-correction algorithms (SVR, LSTM, Bi-LSTM) at Karimun Jawa and Bira, finding a warm bias at both sites, with LSTM best at Karimun Jawa (RMSE 0.207, MAE 0.182, MBE -0.171, a 29.83% RMSE reduction) and Bi-LSTM best at Bira (RMSE 0.218, MAE 0.173, MBE 0.011, a 26.10% RMSE reduction), indicating that correction effectiveness is location-dependent and algorithm choice should rest on local validation.
arXiv

Reflex-Guard cuts prompt-safety filtering to 37.6 ms locally with dense semantic embeddings, reaching 95.9% harmful-prompt recall on 30,568 samples

The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
arXiv

TP-CRIV lets a third party with no white-box or API access verify an AI model's identity using only its ordinary black-box inference interface

The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a setting where the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under fresh, previously undisclosed requirements and network isolation it obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models showing clear same/cross-model separation and finite-chal
arXiv

LiveMathematicianBench builds a research-level math reasoning benchmark from fresh arXiv papers, where the best model reaches only 43.5% and Gemini-3.1-pro-preview falls to 17.6% under substitution-resistant evaluation

The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs, introducing a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism; evaluation shows the benchmark is far from saturated, with the best model Gemini-3.1-pro-preview reaching only 43.5%, substitution-resistant evaluation yielding a top score of 30.6% for GPT-5.4 while Gemini-3.1-pro-preview drops to 17.6% below the 20% random baseline, and a dual-mode protocol showing consistent accuracy gains from proof-sketch access.
arXiv

DMA approximates factor-to-variable messages directly, bounds marginal KL by message KL, and yields a Bayesian neural network inference algorithm without a learning-rate hyperparameter

The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod