arXiv The work builds a computational library of 146 single-walled carbon nanotube (SWCNT)–SWCNT junctions and analyses their magnetotransport with an automated workflow combining molecular dynamics, tight-binding theory, Peierls magnetic coupling, and non-equilibrium Green's functions, followed by machine-learning analysis; it finds that the averaged first transmission-step value is governed primarily by the mean chiral angle of the two nanotubes, that the junction energy gap depends predominantly on the metallic or semiconducting character of the constituent nanotubes, that temperature generally suppresses the averaged transmission while reducing the extracted gap, that a perpendicular magnetic field affects transmission much more strongly than the gap, that signatures of interference-driven t
The work builds a computational library of 146 single-walled carbon nanotube (SWCNT)–SWCNT junctions and analyses their magnetotransport with an automated workflow combining molecular dynamics, tight-binding theory, Peierls magnetic coupling, and non-equilibrium Green's functions, followed by machine-learning analysis; it finds that the averaged first transmission-step value is governed primarily by the mean chiral angle of the two nanotubes, that the junction energy gap depends predominantly on the metallic or semiconducting character of the constituent nanotubes, that temperature generally suppresses the averaged transmission while reducing the extracted gap, that a perpendicular magnetic field affects transmission much more strongly than the gap, that signatures of interference-driven t
The work builds a computational library of 146 single-walled carbon nanotube (SWCNT)–SWCNT junctions and analyses their magnetotransport with an automated workflow combining molecular dynamics, tight-binding theory, Peierls magnetic coupling, and non-equilibrium Green's functions, followed by machine-learning analysis; it finds that the averaged first transmission-step value is governed primarily by the mean chiral angle of the two nanotubes, that the junction energy gap depends predominantly on the metallic or semiconducting character of the constituent nanotubes, that temperature generally suppresses the averaged transmission while reducing the extracted gap, that a perpendicular magnetic field affects transmission much more strongly than the gap, that signatures of interference-driven t
The work builds a computational library of 146 single-walled carbon nanotube (SWCNT)–SWCNT junctions and analyses their magnetotransport with an automated workflow combining molecular dynamics, tight-binding theory, Peierls magnetic coupling, and non-equilibrium Green's functions, followed by machine-learning analysis; it finds that the averaged first transmission-step value is governed primarily by the mean chiral angle of the two nanotubes, that the junction energy gap depends predominantly on the metallic or semiconducting character of the constituent nanotubes, that temperature generally suppresses the averaged transmission while reducing the extracted gap, that a perpendicular magnetic field affects transmission much more strongly than the gap, that signatures of interference-driven t
arXiv The work formalizes audio description (AD) generation as a constrained optimization problem over three coupled decisions—what visual information to describe, when it can be spoken in gaps between dialogue, and how to formulate it to fit the available time—and proposes a hybrid system in which large language models propose and ground visual elements, estimate their narrative salience, and generate compressed realizations, while a mixed-integer linear program jointly selects and schedules descriptions across a scene under temporal constraints; evaluated on REFRAMED, a benchmark for realistic AD of movies, the approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new state of the art on narrative QA and temporally grounded metrics, w
The work formalizes audio description (AD) generation as a constrained optimization problem over three coupled decisions—what visual information to describe, when it can be spoken in gaps between dialogue, and how to formulate it to fit the available time—and proposes a hybrid system in which large language models propose and ground visual elements, estimate their narrative salience, and generate compressed realizations, while a mixed-integer linear program jointly selects and schedules descriptions across a scene under temporal constraints; evaluated on REFRAMED, a benchmark for realistic AD of movies, the approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new state of the art on narrative QA and temporally grounded metrics, w
The work formalizes audio description (AD) generation as a constrained optimization problem over three coupled decisions—what visual information to describe, when it can be spoken in gaps between dialogue, and how to formulate it to fit the available time—and proposes a hybrid system in which large language models propose and ground visual elements, estimate their narrative salience, and generate compressed realizations, while a mixed-integer linear program jointly selects and schedules descriptions across a scene under temporal constraints; evaluated on REFRAMED, a benchmark for realistic AD of movies, the approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new state of the art on narrative QA and temporally grounded metrics, w
The work formalizes audio description (AD) generation as a constrained optimization problem over three coupled decisions—what visual information to describe, when it can be spoken in gaps between dialogue, and how to formulate it to fit the available time—and proposes a hybrid system in which large language models propose and ground visual elements, estimate their narrative salience, and generate compressed realizations, while a mixed-integer linear program jointly selects and schedules descriptions across a scene under temporal constraints; evaluated on REFRAMED, a benchmark for realistic AD of movies, the approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new state of the art on narrative QA and temporally grounded metrics, w
arXiv OncoVision is a privileged-information training framework that uses mammography images and clinical features during training while performing inference from mammographic images alone; built on an attention-based encoder-decoder backbone, it jointly segments four regions of interest (masses, calcifications, axillary findings, and breast tissue) with accuracy exceeding the nnU-Net baseline and predicts ten structured clinical features including BI-RADS category; the authors developed two late-fusion strategies, Independent and Dependent, that integrate imaging, radiomic, and clinical information during training, with radiomic features extracted from predicted masks providing shape, intensity, and texture descriptors that complement the learned CNN representations; in a retrospective multi-re
OncoVision is a privileged-information training framework that uses mammography images and clinical features during training while performing inference from mammographic images alone; built on an attention-based encoder-decoder backbone, it jointly segments four regions of interest (masses, calcifications, axillary findings, and breast tissue) with accuracy exceeding the nnU-Net baseline and predicts ten structured clinical features including BI-RADS category; the authors developed two late-fusion strategies, Independent and Dependent, that integrate imaging, radiomic, and clinical information during training, with radiomic features extracted from predicted masks providing shape, intensity, and texture descriptors that complement the learned CNN representations; in a retrospective multi-re
OncoVision is a privileged-information training framework that uses mammography images and clinical features during training while performing inference from mammographic images alone; built on an attention-based encoder-decoder backbone, it jointly segments four regions of interest (masses, calcifications, axillary findings, and breast tissue) with accuracy exceeding the nnU-Net baseline and predicts ten structured clinical features including BI-RADS category; the authors developed two late-fusion strategies, Independent and Dependent, that integrate imaging, radiomic, and clinical information during training, with radiomic features extracted from predicted masks providing shape, intensity, and texture descriptors that complement the learned CNN representations; in a retrospective multi-re
OncoVision is a privileged-information training framework that uses mammography images and clinical features during training while performing inference from mammographic images alone; built on an attention-based encoder-decoder backbone, it jointly segments four regions of interest (masses, calcifications, axillary findings, and breast tissue) with accuracy exceeding the nnU-Net baseline and predicts ten structured clinical features including BI-RADS category; the authors developed two late-fusion strategies, Independent and Dependent, that integrate imaging, radiomic, and clinical information during training, with radiomic features extracted from predicted masks providing shape, intensity, and texture descriptors that complement the learned CNN representations; in a retrospective multi-re
arXiv SkillGym introduces a framework that transforms human-written agent skills into executable, verifiable training environments: its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions; it constructs and releases 2,756 environments across 12 categories and collects 8,364 successful trajectories (averaging 49 tool calls and over 60k logged text tokens), supporting supervised fine-tuning on verified workflows and reinforcement learning with outcome-based rewards; under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.
SkillGym introduces a framework that transforms human-written agent skills into executable, verifiable training environments: its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions; it constructs and releases 2,756 environments across 12 categories and collects 8,364 successful trajectories (averaging 49 tool calls and over 60k logged text tokens), supporting supervised fine-tuning on verified workflows and reinforcement learning with outcome-based rewards; under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.
SkillGym introduces a framework that transforms human-written agent skills into executable, verifiable training environments: its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions; it constructs and releases 2,756 environments across 12 categories and collects 8,364 successful trajectories (averaging 49 tool calls and over 60k logged text tokens), supporting supervised fine-tuning on verified workflows and reinforcement learning with outcome-based rewards; under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.
SkillGym introduces a framework that transforms human-written agent skills into executable, verifiable training environments: its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions; it constructs and releases 2,756 environments across 12 categories and collects 8,364 successful trajectories (averaging 49 tool calls and over 60k logged text tokens), supporting supervised fine-tuning on verified workflows and reinforcement learning with outcome-based rewards; under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.
arXiv STAM-ASR is a lightweight framework that, without relying on an external diarization system and without explicit speech separation, learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features to give explicit who-and-when cues that modulate the AudioLLM's semantic representation, and maintains fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
STAM-ASR is a lightweight framework that, without relying on an external diarization system and without explicit speech separation, learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features to give explicit who-and-when cues that modulate the AudioLLM's semantic representation, and maintains fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
STAM-ASR is a lightweight framework that, without relying on an external diarization system and without explicit speech separation, learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features to give explicit who-and-when cues that modulate the AudioLLM's semantic representation, and maintains fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
STAM-ASR is a lightweight framework that, without relying on an external diarization system and without explicit speech separation, learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features to give explicit who-and-when cues that modulate the AudioLLM's semantic representation, and maintains fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
arXiv TIDE is a serving-engine-native framework that reuses the target model's intermediate hidden states produced during inference as training signals for online draft adaptation, avoiding additional target model computation and serving-time overhead, and combines this with adaptive runtime control that activates speculation and draft training only when beneficial and with mapping of inference and training onto appropriate GPU classes; across diverse real-world workloads TIDE achieves up to 1.66x throughput over no-speculation baselines, recovers performance on misaligned workloads where static draft models degrade throughput, reduces training time by up to 3.02x and storage requirements by 24x compared to existing draft training approaches, and improves system throughput by up to 1.
TIDE is a serving-engine-native framework that reuses the target model's intermediate hidden states produced during inference as training signals for online draft adaptation, avoiding additional target model computation and serving-time overhead, and combines this with adaptive runtime control that activates speculation and draft training only when beneficial and with mapping of inference and training onto appropriate GPU classes; across diverse real-world workloads TIDE achieves up to 1.66x throughput over no-speculation baselines, recovers performance on misaligned workloads where static draft models degrade throughput, reduces training time by up to 3.02x and storage requirements by 24x compared to existing draft training approaches, and improves system throughput by up to 1.
TIDE is a serving-engine-native framework that reuses the target model's intermediate hidden states produced during inference as training signals for online draft adaptation, avoiding additional target model computation and serving-time overhead, and combines this with adaptive runtime control that activates speculation and draft training only when beneficial and with mapping of inference and training onto appropriate GPU classes; across diverse real-world workloads TIDE achieves up to 1.66x throughput over no-speculation baselines, recovers performance on misaligned workloads where static draft models degrade throughput, reduces training time by up to 3.02x and storage requirements by 24x compared to existing draft training approaches, and improves system throughput by up to 1.
TIDE is a serving-engine-native framework that reuses the target model's intermediate hidden states produced during inference as training signals for online draft adaptation, avoiding additional target model computation and serving-time overhead, and combines this with adaptive runtime control that activates speculation and draft training only when beneficial and with mapping of inference and training onto appropriate GPU classes; across diverse real-world workloads TIDE achieves up to 1.66x throughput over no-speculation baselines, recovers performance on misaligned workloads where static draft models degrade throughput, reduces training time by up to 3.02x and storage requirements by 24x compared to existing draft training approaches, and improves system throughput by up to 1.
arXiv The work proposes SEEK (Skill-routed Evaluation with Evolvable Knowledge), which externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution; a two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining; experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis, and SEEK has been deployed at Kuaishou, a platform with over 400 million daily activ
The work proposes SEEK (Skill-routed Evaluation with Evolvable Knowledge), which externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution; a two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining; experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis, and SEEK has been deployed at Kuaishou, a platform with over 400 million daily activ
The work proposes SEEK (Skill-routed Evaluation with Evolvable Knowledge), which externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution; a two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining; experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis, and SEEK has been deployed at Kuaishou, a platform with over 400 million daily activ
The work proposes SEEK (Skill-routed Evaluation with Evolvable Knowledge), which externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution; a two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining; experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis, and SEEK has been deployed at Kuaishou, a platform with over 400 million daily activ
arXiv The work builds Qwen-Planner-Agent within a closed-loop AI-for-AI framework that links data production, model training, and deployment through a shared action-feedback-verification contract: AI for Data builds a human-gated agentic data flywheel, AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning and introduces CARE to reduce reasoning and tool-use costs, and AI drives model-harness co-evolution; the agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination, with further gains on non-mobile agentic benchmarks while largely preserving general capabilities.
The work builds Qwen-Planner-Agent within a closed-loop AI-for-AI framework that links data production, model training, and deployment through a shared action-feedback-verification contract: AI for Data builds a human-gated agentic data flywheel, AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning and introduces CARE to reduce reasoning and tool-use costs, and AI drives model-harness co-evolution; the agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination, with further gains on non-mobile agentic benchmarks while largely preserving general capabilities.
The work builds Qwen-Planner-Agent within a closed-loop AI-for-AI framework that links data production, model training, and deployment through a shared action-feedback-verification contract: AI for Data builds a human-gated agentic data flywheel, AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning and introduces CARE to reduce reasoning and tool-use costs, and AI drives model-harness co-evolution; the agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination, with further gains on non-mobile agentic benchmarks while largely preserving general capabilities.
The work builds Qwen-Planner-Agent within a closed-loop AI-for-AI framework that links data production, model training, and deployment through a shared action-feedback-verification contract: AI for Data builds a human-gated agentic data flywheel, AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning and introduces CARE to reduce reasoning and tool-use costs, and AI drives model-harness co-evolution; the agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination, with further gains on non-mobile agentic benchmarks while largely preserving general capabilities.
arXiv The work introduces Unite-Audio, which jointly learns continuous audio representations and latent flow matching for text-to-audio generation; by coupling reconstruction with self-supervised generative prediction it lets the generative objective directly shape the latent space, and it applies Flow-GRPO post-training to improve text-conditioned generation; experiments show competitive text-to-audio performance with a compact latent flow model, and ablation studies confirm the benefit of jointly learning the audio representation and the generative model.
The work introduces Unite-Audio, which jointly learns continuous audio representations and latent flow matching for text-to-audio generation; by coupling reconstruction with self-supervised generative prediction it lets the generative objective directly shape the latent space, and it applies Flow-GRPO post-training to improve text-conditioned generation; experiments show competitive text-to-audio performance with a compact latent flow model, and ablation studies confirm the benefit of jointly learning the audio representation and the generative model.
The work introduces Unite-Audio, which jointly learns continuous audio representations and latent flow matching for text-to-audio generation; by coupling reconstruction with self-supervised generative prediction it lets the generative objective directly shape the latent space, and it applies Flow-GRPO post-training to improve text-conditioned generation; experiments show competitive text-to-audio performance with a compact latent flow model, and ablation studies confirm the benefit of jointly learning the audio representation and the generative model.
The work introduces Unite-Audio, which jointly learns continuous audio representations and latent flow matching for text-to-audio generation; by coupling reconstruction with self-supervised generative prediction it lets the generative objective directly shape the latent space, and it applies Flow-GRPO post-training to improve text-conditioned generation; experiments show competitive text-to-audio performance with a compact latent flow model, and ablation studies confirm the benefit of jointly learning the audio representation and the generative model.
arXiv The work introduces Self-Play Pretraining with Zero Data as an initial proof-of-concept: starting from random initialization, a generator proposes programs interpreted by a universal Turing machine to produce byte sequences while a learner autoregressively predicts those byte sequences with standard cross-entropy, and the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum; because neither generator nor learner is trained on natural data, zero-shot performance on natural data serves as a clean test of transfer, and across several natural datasets zero-shot loss exhibits predictable scaling in compute, with the models also exhibiting in-context learning and discovering recognizable mathematical
The work introduces Self-Play Pretraining with Zero Data as an initial proof-of-concept: starting from random initialization, a generator proposes programs interpreted by a universal Turing machine to produce byte sequences while a learner autoregressively predicts those byte sequences with standard cross-entropy, and the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum; because neither generator nor learner is trained on natural data, zero-shot performance on natural data serves as a clean test of transfer, and across several natural datasets zero-shot loss exhibits predictable scaling in compute, with the models also exhibiting in-context learning and discovering recognizable mathematical
The work introduces Self-Play Pretraining with Zero Data as an initial proof-of-concept: starting from random initialization, a generator proposes programs interpreted by a universal Turing machine to produce byte sequences while a learner autoregressively predicts those byte sequences with standard cross-entropy, and the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum; because neither generator nor learner is trained on natural data, zero-shot performance on natural data serves as a clean test of transfer, and across several natural datasets zero-shot loss exhibits predictable scaling in compute, with the models also exhibiting in-context learning and discovering recognizable mathematical
The work introduces Self-Play Pretraining with Zero Data as an initial proof-of-concept: starting from random initialization, a generator proposes programs interpreted by a universal Turing machine to produce byte sequences while a learner autoregressively predicts those byte sequences with standard cross-entropy, and the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum; because neither generator nor learner is trained on natural data, zero-shot performance on natural data serves as a clean test of transfer, and across several natural datasets zero-shot loss exhibits predictable scaling in compute, with the models also exhibiting in-context learning and discovering recognizable mathematical
arXiv The work proposes MILO, a compression framework that exploits low-rank redundancy in many-shot contexts by compressing the KV cache at block granularity and dynamically allocating rank budgets based on information entropy, achieving up to 50% KV cache memory reduction and 1.8x throughput improvement on Qwen2.5 models with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.
The work proposes MILO, a compression framework that exploits low-rank redundancy in many-shot contexts by compressing the KV cache at block granularity and dynamically allocating rank budgets based on information entropy, achieving up to 50% KV cache memory reduction and 1.8x throughput improvement on Qwen2.5 models with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.
The work proposes MILO, a compression framework that exploits low-rank redundancy in many-shot contexts by compressing the KV cache at block granularity and dynamically allocating rank budgets based on information entropy, achieving up to 50% KV cache memory reduction and 1.8x throughput improvement on Qwen2.5 models with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.
The work proposes MILO, a compression framework that exploits low-rank redundancy in many-shot contexts by compressing the KV cache at block granularity and dynamically allocating rank budgets based on information entropy, achieving up to 50% KV cache memory reduction and 1.8x throughput improvement on Qwen2.5 models with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.
arXiv The work releases a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI with no human-authored code or tests, together with two code-provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability; applying these to the dataset, it finds that user coding agent CLI instructions differ in kind from IDE-chat instructions with a greater focus on comprehension, planning and consultation, that code development is mainly proactive, that 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, and that roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.
The work releases a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI with no human-authored code or tests, together with two code-provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability; applying these to the dataset, it finds that user coding agent CLI instructions differ in kind from IDE-chat instructions with a greater focus on comprehension, planning and consultation, that code development is mainly proactive, that 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, and that roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.
The work releases a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI with no human-authored code or tests, together with two code-provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability; applying these to the dataset, it finds that user coding agent CLI instructions differ in kind from IDE-chat instructions with a greater focus on comprehension, planning and consultation, that code development is mainly proactive, that 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, and that roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.
The work releases a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI with no human-authored code or tests, together with two code-provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability; applying these to the dataset, it finds that user coding agent CLI instructions differ in kind from IDE-chat instructions with a greater focus on comprehension, planning and consultation, that code development is mainly proactive, that 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, and that roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.
arXiv The work introduces WebArxiv, a static-snapshot benchmark built on arXiv with 510 time-invariant tasks, each with a unique deterministic ground truth, for evaluating multimodal web agents on scholarly tasks such as multi-constraint paper retrieval, fine-grained content extraction, and cross-paper comparison; evaluations of a range of foundation-model-based web agents show the benchmark remains challenging, behavioral analysis reveals that agents over-rely on fixed interaction histories, causing incomplete or repetitive reasoning, and the authors therefore equip agents with a lightweight dynamic-memory mechanism for adaptive retrieval and reasoning over relevant context.
The work introduces WebArxiv, a static-snapshot benchmark built on arXiv with 510 time-invariant tasks, each with a unique deterministic ground truth, for evaluating multimodal web agents on scholarly tasks such as multi-constraint paper retrieval, fine-grained content extraction, and cross-paper comparison; evaluations of a range of foundation-model-based web agents show the benchmark remains challenging, behavioral analysis reveals that agents over-rely on fixed interaction histories, causing incomplete or repetitive reasoning, and the authors therefore equip agents with a lightweight dynamic-memory mechanism for adaptive retrieval and reasoning over relevant context.
The work introduces WebArxiv, a static-snapshot benchmark built on arXiv with 510 time-invariant tasks, each with a unique deterministic ground truth, for evaluating multimodal web agents on scholarly tasks such as multi-constraint paper retrieval, fine-grained content extraction, and cross-paper comparison; evaluations of a range of foundation-model-based web agents show the benchmark remains challenging, behavioral analysis reveals that agents over-rely on fixed interaction histories, causing incomplete or repetitive reasoning, and the authors therefore equip agents with a lightweight dynamic-memory mechanism for adaptive retrieval and reasoning over relevant context.
The work introduces WebArxiv, a static-snapshot benchmark built on arXiv with 510 time-invariant tasks, each with a unique deterministic ground truth, for evaluating multimodal web agents on scholarly tasks such as multi-constraint paper retrieval, fine-grained content extraction, and cross-paper comparison; evaluations of a range of foundation-model-based web agents show the benchmark remains challenging, behavioral analysis reveals that agents over-rely on fixed interaction histories, causing incomplete or repetitive reasoning, and the authors therefore equip agents with a lightweight dynamic-memory mechanism for adaptive retrieval and reasoning over relevant context.
arXiv The work identifies an evaluation practice it calls CoT-prefix scoring, in which a reasoning cue is appended to the prompt but answer-label logits are read before the model generates any rationale, and reports that on ScienceQA Qwen2.5-VL-7B falls from 80.76% to 45.48%, that 93.54% of CoT-prefix predictions select the first slot across five option-content permutations, and that condition-matched linear probes recover 78.94% from the same hidden states while free generation restores 75.24%, with vocabulary and layer diagnostics showing probability mass shifting toward continuation tokens while answer information remains linearly accessible in late layers, indicating an evaluation-interface mismatch rather than missing model knowledge.
The work identifies an evaluation practice it calls CoT-prefix scoring, in which a reasoning cue is appended to the prompt but answer-label logits are read before the model generates any rationale, and reports that on ScienceQA Qwen2.5-VL-7B falls from 80.76% to 45.48%, that 93.54% of CoT-prefix predictions select the first slot across five option-content permutations, and that condition-matched linear probes recover 78.94% from the same hidden states while free generation restores 75.24%, with vocabulary and layer diagnostics showing probability mass shifting toward continuation tokens while answer information remains linearly accessible in late layers, indicating an evaluation-interface mismatch rather than missing model knowledge.
The work identifies an evaluation practice it calls CoT-prefix scoring, in which a reasoning cue is appended to the prompt but answer-label logits are read before the model generates any rationale, and reports that on ScienceQA Qwen2.5-VL-7B falls from 80.76% to 45.48%, that 93.54% of CoT-prefix predictions select the first slot across five option-content permutations, and that condition-matched linear probes recover 78.94% from the same hidden states while free generation restores 75.24%, with vocabulary and layer diagnostics showing probability mass shifting toward continuation tokens while answer information remains linearly accessible in late layers, indicating an evaluation-interface mismatch rather than missing model knowledge.
The work identifies an evaluation practice it calls CoT-prefix scoring, in which a reasoning cue is appended to the prompt but answer-label logits are read before the model generates any rationale, and reports that on ScienceQA Qwen2.5-VL-7B falls from 80.76% to 45.48%, that 93.54% of CoT-prefix predictions select the first slot across five option-content permutations, and that condition-matched linear probes recover 78.94% from the same hidden states while free generation restores 75.24%, with vocabulary and layer diagnostics showing probability mass shifting toward continuation tokens while answer information remains linearly accessible in late layers, indicating an evaluation-interface mismatch rather than missing model knowledge.
arXiv This work audits a multilingual affective generation benchmark in which eight instruction-tuned LLMs produced emoji summaries for 17,100 Bangla, English and Hindi sentences with 6,960 human judgements, finding that when annotators are treated as a random rather than a fixed factor no system differs significantly from any other (F(7,14)=0.59, p=0.76) although the conventional analysis declares 19 of 28 pairwise differences significant; annotator identity explains far more rating variance than system identity and removing any single annotator changes the winning system; the ordering that emerges tracks output length, with mean emoji count explaining 78.
This work audits a multilingual affective generation benchmark in which eight instruction-tuned LLMs produced emoji summaries for 17,100 Bangla, English and Hindi sentences with 6,960 human judgements, finding that when annotators are treated as a random rather than a fixed factor no system differs significantly from any other (F(7,14)=0.59, p=0.76) although the conventional analysis declares 19 of 28 pairwise differences significant; annotator identity explains far more rating variance than system identity and removing any single annotator changes the winning system; the ordering that emerges tracks output length, with mean emoji count explaining 78.
This work audits a multilingual affective generation benchmark in which eight instruction-tuned LLMs produced emoji summaries for 17,100 Bangla, English and Hindi sentences with 6,960 human judgements, finding that when annotators are treated as a random rather than a fixed factor no system differs significantly from any other (F(7,14)=0.59, p=0.76) although the conventional analysis declares 19 of 28 pairwise differences significant; annotator identity explains far more rating variance than system identity and removing any single annotator changes the winning system; the ordering that emerges tracks output length, with mean emoji count explaining 78.
This work audits a multilingual affective generation benchmark in which eight instruction-tuned LLMs produced emoji summaries for 17,100 Bangla, English and Hindi sentences with 6,960 human judgements, finding that when annotators are treated as a random rather than a fixed factor no system differs significantly from any other (F(7,14)=0.59, p=0.76) although the conventional analysis declares 19 of 28 pairwise differences significant; annotator identity explains far more rating variance than system identity and removing any single annotator changes the winning system; the ordering that emerges tracks output length, with mean emoji count explaining 78.
arXiv The work develops a task-substitution framework for Digital Governance Frameworks (DGF) that treats each governance gate as an executable contract, requiring sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted; it derives a residual-work threshold showing why automating most cases can still increase labor, and measures models on DGF-Bench (300 synthetic projects, 899 evaluable model-project runs): Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%, with complete-route success of 76.92%, 42.33%, and 24.
The work develops a task-substitution framework for Digital Governance Frameworks (DGF) that treats each governance gate as an executable contract, requiring sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted; it derives a residual-work threshold showing why automating most cases can still increase labor, and measures models on DGF-Bench (300 synthetic projects, 899 evaluable model-project runs): Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%, with complete-route success of 76.92%, 42.33%, and 24.
The work develops a task-substitution framework for Digital Governance Frameworks (DGF) that treats each governance gate as an executable contract, requiring sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted; it derives a residual-work threshold showing why automating most cases can still increase labor, and measures models on DGF-Bench (300 synthetic projects, 899 evaluable model-project runs): Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%, with complete-route success of 76.92%, 42.33%, and 24.
The work develops a task-substitution framework for Digital Governance Frameworks (DGF) that treats each governance gate as an executable contract, requiring sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted; it derives a residual-work threshold showing why automating most cases can still increase labor, and measures models on DGF-Bench (300 synthetic projects, 899 evaluable model-project runs): Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%, with complete-route success of 76.92%, 42.33%, and 24.
arXiv The paper argues that patient-reported pain location is sometimes diagnostically decisive and sometimes nearly uninformative because it contains three epistemically distinct failures—anatomical multiplexing as a non-identifiable inverse problem, delocalized amplification (clinically central sensitization or nociplastic pain) as a change of generative model, and referred and atypical displacement hypothesized as a systematic, person-dependent shift—which are one Bayesian inference problem failing at the likelihood, the model class, and the group-conditional prior, with a fourth node at the report itself; on this basis the author finds that the published "high-utility" accuracy band leans on overstated specificity, so the gradient is real but flatter than drawn, and attributes the finding th
The paper argues that patient-reported pain location is sometimes diagnostically decisive and sometimes nearly uninformative because it contains three epistemically distinct failures—anatomical multiplexing as a non-identifiable inverse problem, delocalized amplification (clinically central sensitization or nociplastic pain) as a change of generative model, and referred and atypical displacement hypothesized as a systematic, person-dependent shift—which are one Bayesian inference problem failing at the likelihood, the model class, and the group-conditional prior, with a fourth node at the report itself; on this basis the author finds that the published "high-utility" accuracy band leans on overstated specificity, so the gradient is real but flatter than drawn, and attributes the finding th
The paper argues that patient-reported pain location is sometimes diagnostically decisive and sometimes nearly uninformative because it contains three epistemically distinct failures—anatomical multiplexing as a non-identifiable inverse problem, delocalized amplification (clinically central sensitization or nociplastic pain) as a change of generative model, and referred and atypical displacement hypothesized as a systematic, person-dependent shift—which are one Bayesian inference problem failing at the likelihood, the model class, and the group-conditional prior, with a fourth node at the report itself; on this basis the author finds that the published "high-utility" accuracy band leans on overstated specificity, so the gradient is real but flatter than drawn, and attributes the finding th
The paper argues that patient-reported pain location is sometimes diagnostically decisive and sometimes nearly uninformative because it contains three epistemically distinct failures—anatomical multiplexing as a non-identifiable inverse problem, delocalized amplification (clinically central sensitization or nociplastic pain) as a change of generative model, and referred and atypical displacement hypothesized as a systematic, person-dependent shift—which are one Bayesian inference problem failing at the likelihood, the model class, and the group-conditional prior, with a fourth node at the report itself; on this basis the author finds that the published "high-utility" accuracy band leans on overstated specificity, so the gradient is real but flatter than drawn, and attributes the finding th
arXiv The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
arXiv The authors propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms (STFT, continuous wavelet transform, discrete wavelet transform), and extend the virtual inspection layer from the STFT to both wavelet transforms so that LRP can generate explanations in these representations; on a synthetic dataset XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines, and across two real-world datasets it produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria.
The authors propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms (STFT, continuous wavelet transform, discrete wavelet transform), and extend the virtual inspection layer from the STFT to both wavelet transforms so that LRP can generate explanations in these representations; on a synthetic dataset XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines, and across two real-world datasets it produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria.
The authors propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms (STFT, continuous wavelet transform, discrete wavelet transform), and extend the virtual inspection layer from the STFT to both wavelet transforms so that LRP can generate explanations in these representations; on a synthetic dataset XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines, and across two real-world datasets it produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria.
The authors propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms (STFT, continuous wavelet transform, discrete wavelet transform), and extend the virtual inspection layer from the STFT to both wavelet transforms so that LRP can generate explanations in these representations; on a synthetic dataset XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines, and across two real-world datasets it produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria.
arXiv The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a third-party setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under these constraints the framework obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models demonstrating clear same/cross-model separation
The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a third-party setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under these constraints the framework obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models demonstrating clear same/cross-model separation
The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a third-party setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under these constraints the framework obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models demonstrating clear same/cross-model separation
The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a third-party setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under these constraints the framework obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models demonstrating clear same/cross-model separation