Skip to main content

Search

“All disciplines” · 1495 results

Page 35 · showing 20
arXiv

A 146-junction single-walled carbon nanotube library shows mean chiral angle sets the averaged transmission step while metallic or semiconducting character sets the junction gap

The work builds a computational library of 146 single-walled carbon nanotube (SWCNT)–SWCNT junctions and analyses their magnetotransport with an automated workflow combining molecular dynamics, tight-binding theory, Peierls magnetic coupling, and non-equilibrium Green's functions, followed by machine-learning analysis; it finds that the averaged first transmission-step value is governed primarily by the mean chiral angle of the two nanotubes, that the junction energy gap depends predominantly on the metallic or semiconducting character of the constituent nanotubes, that temperature generally suppresses the averaged transmission while reducing the extracted gap, that a perpendicular magnetic field affects transmission much more strongly than the gap, that signatures of interference-driven t
arXiv

Reframing audio description as constrained optimization over what, when, and how lets a hybrid LLM-plus-MILP system set new state of the art on REFRAMED's narrative QA and temporal metrics

The work formalizes audio description (AD) generation as a constrained optimization problem over three coupled decisions—what visual information to describe, when it can be spoken in gaps between dialogue, and how to formulate it to fit the available time—and proposes a hybrid system in which large language models propose and ground visual elements, estimate their narrative salience, and generate compressed realizations, while a mixed-integer linear program jointly selects and schedules descriptions across a scene under temporal constraints; evaluated on REFRAMED, a benchmark for realistic AD of movies, the approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new state of the art on narrative QA and temporally grounded metrics, w
arXiv

OncoVision's attention-driven multimodal training framework cut reading time by up to 61% and raised diagnostic confidence in a paired six-radiologist evaluation

OncoVision is a privileged-information training framework that uses mammography images and clinical features during training while performing inference from mammographic images alone; built on an attention-based encoder-decoder backbone, it jointly segments four regions of interest (masses, calcifications, axillary findings, and breast tissue) with accuracy exceeding the nnU-Net baseline and predicts ten structured clinical features including BI-RADS category; the authors developed two late-fusion strategies, Independent and Dependent, that integrate imaging, radiomic, and clinical information during training, with radiomic features extracted from predicted masks providing shape, intensity, and texture descriptors that complement the learned CNN representations; in a retrospective multi-re
arXiv

SkillGym turns human skills into verifiable training environments, letting a 35B model reach 51.47% on SkillsBench and exceed reported scores of Claude Sonnet 4.6 and others

SkillGym introduces a framework that transforms human-written agent skills into executable, verifiable training environments: its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions; it constructs and releases 2,756 environments across 12 categories and collects 8,364 successful trajectories (averaging 49 tool calls and over 60k logged text tokens), supporting supervised fine-tuning on verified workflows and reinforcement learning with outcome-based rewards; under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.
arXiv

STAM-ASR extends a pretrained AudioLLM with speaker-temporal anchoring and memory for multi-speaker ASR, evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1

STAM-ASR is a lightweight framework that, without relying on an external diarization system and without explicit speech separation, learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features to give explicit who-and-when cues that modulate the AudioLLM's semantic representation, and maintains fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
arXiv

TIDE reuses inference-time hidden states to adapt draft models online, reaching up to 1.66x throughput and recovering performance when static drafts degrade it

TIDE is a serving-engine-native framework that reuses the target model's intermediate hidden states produced during inference as training signals for online draft adaptation, avoiding additional target model computation and serving-time overhead, and combines this with adaptive runtime control that activates speculation and draft training only when beneficial and with mapping of inference and training onto appropriate GPU classes; across diverse real-world workloads TIDE achieves up to 1.66x throughput over no-speculation baselines, recovers performance on misaligned workloads where static draft models degrade throughput, reduces training time by up to 3.02x and storage requirements by 24x compared to existing draft training approaches, and improves system throughput by up to 1.
arXiv

SEEK externalizes search evaluation criteria into a routable skill bank, improving listwise quality evaluation and attribution diagnosis in Kuaishou short-video search

The work proposes SEEK (Skill-routed Evaluation with Evolvable Knowledge), which externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution; a two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining; experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis, and SEEK has been deployed at Kuaishou, a platform with over 400 million daily activ
arXiv

Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model in tool use, memory, skills, and sub-agent coordination

The work builds Qwen-Planner-Agent within a closed-loop AI-for-AI framework that links data production, model training, and deployment through a shared action-feedback-verification contract: AI for Data builds a human-gated agentic data flywheel, AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning and introduces CARE to reduce reasoning and tool-use costs, and AI drives model-harness co-evolution; the agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination, with further gains on non-mobile agentic benchmarks while largely preserving general capabilities.
arXiv

Unite-Audio jointly trains audio representation learning with latent flow matching, reaching competitive text-to-audio generation with a compact model

The work introduces Unite-Audio, which jointly learns continuous audio representations and latent flow matching for text-to-audio generation; by coupling reconstruction with self-supervised generative prediction it lets the generative objective directly shape the latent space, and it applies Flow-GRPO post-training to improve text-conditioned generation; experiments show competitive text-to-audio performance with a compact latent flow model, and ablation studies confirm the benefit of jointly learning the audio representation and the generative model.
arXiv

Zero-Data Self-Play Pretraining: a generator and a learner trained in tandem from random initialization show zero-shot loss scaling predictably with self-play compute across several natural datasets

The work introduces Self-Play Pretraining with Zero Data as an initial proof-of-concept: starting from random initialization, a generator proposes programs interpreted by a universal Turing machine to produce byte sequences while a learner autoregressively predicts those byte sequences with standard cross-entropy, and the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum; because neither generator nor learner is trained on natural data, zero-shot performance on natural data serves as a clean test of transfer, and across several natural datasets zero-shot loss exhibits predictable scaling in compute, with the models also exhibiting in-context learning and discovering recognizable mathematical
arXiv

MILO cuts many-shot context KV cache by up to 50% via block-wise low-rank compression, lifting Qwen2.5 throughput 1.8x with near-flat classification and reasoning performance

The work proposes MILO, a compression framework that exploits low-rank redundancy in many-shot contexts by compressing the KV cache at block granularity and dynamically allocating rank budgets based on information entropy, achieving up to 50% KV cache memory reduction and 1.8x throughput improvement on Qwen2.5 models with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.
arXiv

Full development history of a wholly AI-authored codebase released: 14.3% of AI code-generation events in a 21,000-line Python tool contained real errors, and roughly 1 in 4-5 interactive responses contained factual errors

The work releases a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI with no human-authored code or tests, together with two code-provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability; applying these to the dataset, it finds that user coding agent CLI instructions differ in kind from IDE-chat instructions with a greater focus on comprehension, planning and consultation, that code development is mainly proactive, that 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, and that roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.
arXiv

WebArxiv evaluates web agents on 510 static-snapshot tasks, finds reliance on fixed interaction histories, and adds a dynamic-memory mechanism

The work introduces WebArxiv, a static-snapshot benchmark built on arXiv with 510 time-invariant tasks, each with a unique deterministic ground truth, for evaluating multimodal web agents on scholarly tasks such as multi-constraint paper retrieval, fine-grained content extraction, and cross-paper comparison; evaluations of a range of foundation-model-based web agents show the benchmark remains challenging, behavioral analysis reveals that agents over-rely on fixed interaction histories, causing incomplete or repetitive reasoning, and the authors therefore equip agents with a lightweight dynamic-memory mechanism for adaptive retrieval and reasoning over relevant context.
arXiv

Reading answer-label logits under a chain-of-thought cue drops Qwen2.5-VL-7B on ScienceQA from 80.76% to 45.48%, with 93.54% of predictions landing in the first option slot

The work identifies an evaluation practice it calls CoT-prefix scoring, in which a reasoning cue is appended to the prompt but answer-label logits are read before the model generates any rationale, and reports that on ScienceQA Qwen2.5-VL-7B falls from 80.76% to 45.48%, that 93.54% of CoT-prefix predictions select the first slot across five option-content permutations, and that condition-matched linear probes recover 78.94% from the same hidden states while free generation restores 75.24%, with vocabulary and layer diagnostics showing probability mass shifting toward continuation tokens while answer information remains linearly accessible in late layers, indicating an evaluation-interface mismatch rather than missing model knowledge.
arXiv

Audit finds a multilingual affective generation benchmark's headline conclusions stem from the measurement instrument, not system differences, and proposes an emoji-affect decodability probe

This work audits a multilingual affective generation benchmark in which eight instruction-tuned LLMs produced emoji summaries for 17,100 Bangla, English and Hindi sentences with 6,960 human judgements, finding that when annotators are treated as a random rather than a fixed factor no system differs significantly from any other (F(7,14)=0.59, p=0.76) although the conventional analysis declares 19 of 28 pairwise differences significant; annotator identity explains far more rating variance than system identity and removing any single annotator changes the winning system; the ordering that emerges tracks output length, with mean emoji count explaining 78.
arXiv

Handing governance review gates to agents: top model reaches 94.98% strict gate success on DGF-Bench but only 76.92% complete-route success

The work develops a task-substitution framework for Digital Governance Frameworks (DGF) that treats each governance gate as an executable contract, requiring sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted; it derives a residual-work threshold showing why automating most cases can still increase labor, and measures models on DGF-Bench (300 synthetic projects, 899 evaluable model-project runs): Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%, with complete-route success of 76.92%, 42.33%, and 24.
arXiv

Pain location's diagnostic value is split into three failures—anatomical multiplexing, central amplification, and person-dependent displacement—not one gradient

The paper argues that patient-reported pain location is sometimes diagnostically decisive and sometimes nearly uninformative because it contains three epistemically distinct failures—anatomical multiplexing as a non-identifiable inverse problem, delocalized amplification (clinically central sensitization or nociplastic pain) as a change of generative model, and referred and atypical displacement hypothesized as a systematic, person-dependent shift—which are one Bayesian inference problem failing at the likelihood, the model class, and the group-conditional prior, with a fourth node at the report itself; on this basis the author finds that the published "high-utility" accuracy band leans on overstated specificity, so the gradient is real but flatter than drawn, and attributes the finding th
arXiv

Reflex-Guard achieves 95.9% harmful-prompt recall at 37.6 ms locally, faster than Llama Guard 2 and SafeDecoding

The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
arXiv

XACT learns sparse attribution masks over invertible time-frequency transform coefficients, highlighting fewer spurious features than baselines on synthetic data and yielding sparse, structured explanations on two real-world datasets

The authors propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms (STFT, continuous wavelet transform, discrete wavelet transform), and extend the virtual inspection layer from the STFT to both wavelet transforms so that LRP can generate explanations in these representations; on a synthetic dataset XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines, and across two real-world datasets it produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria.
arXiv

TP-CRIV proposes a third-party challenge-response identity verification framework, achieving separable same-model and cross-model verification on ten ImageNet-pretrained TorchVision models

The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a third-party setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under these constraints the framework obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models demonstrating clear same/cross-model separation