arXiv The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
arXiv The authors introduce SciMuse, which generates personalized research ideas using a knowledge graph of 58 million papers and a large language model, and they conducted a large-scale evaluation in which more than 100 research group leaders spanning the natural sciences to the humanities rated over 4,400 personalized ideas by level of interest; expert ratings were modest (mean 2.40 on a 5-point scale, most common rating 1) while 24.9% of ideas were rated 4 or 5, supplying knowledge-graph-selected concept pairs did not improve expert-rated interest over a titles-only GPT baseline, high-citation-predicted pairs even showed a weak tendency (1.
The authors introduce SciMuse, which generates personalized research ideas using a knowledge graph of 58 million papers and a large language model, and they conducted a large-scale evaluation in which more than 100 research group leaders spanning the natural sciences to the humanities rated over 4,400 personalized ideas by level of interest; expert ratings were modest (mean 2.40 on a 5-point scale, most common rating 1) while 24.9% of ideas were rated 4 or 5, supplying knowledge-graph-selected concept pairs did not improve expert-rated interest over a titles-only GPT baseline, high-citation-predicted pairs even showed a weak tendency (1.
The authors introduce SciMuse, which generates personalized research ideas using a knowledge graph of 58 million papers and a large language model, and they conducted a large-scale evaluation in which more than 100 research group leaders spanning the natural sciences to the humanities rated over 4,400 personalized ideas by level of interest; expert ratings were modest (mean 2.40 on a 5-point scale, most common rating 1) while 24.9% of ideas were rated 4 or 5, supplying knowledge-graph-selected concept pairs did not improve expert-rated interest over a titles-only GPT baseline, high-citation-predicted pairs even showed a weak tendency (1.
The authors introduce SciMuse, which generates personalized research ideas using a knowledge graph of 58 million papers and a large language model, and they conducted a large-scale evaluation in which more than 100 research group leaders spanning the natural sciences to the humanities rated over 4,400 personalized ideas by level of interest; expert ratings were modest (mean 2.40 on a 5-point scale, most common rating 1) while 24.9% of ideas were rated 4 or 5, supplying knowledge-graph-selected concept pairs did not improve expert-rated interest over a titles-only GPT baseline, high-citation-predicted pairs even showed a weak tendency (1.
arXiv MorphIK is a flow-matching model that encodes a robot's morphology together with the target pose using a transformer and conditions a flow-matching head on that encoding to generate joint poses from noise, thereby solving inverse kinematics for revolute-joint kinematic chains never seen during training; trained only on purely synthetic data from procedurally generated robots, it reaches about 5 cm precision on unseen real-world robots with 6 to 9 Degrees of Freedom, serves as a prior that reduces error to less than 1 cm after a single step of Damped Least Squares optimization and to sub-1 mm after 3 steps in most cases, and can efficiently sample the null space to yield varied configurations for the same pose.
MorphIK is a flow-matching model that encodes a robot's morphology together with the target pose using a transformer and conditions a flow-matching head on that encoding to generate joint poses from noise, thereby solving inverse kinematics for revolute-joint kinematic chains never seen during training; trained only on purely synthetic data from procedurally generated robots, it reaches about 5 cm precision on unseen real-world robots with 6 to 9 Degrees of Freedom, serves as a prior that reduces error to less than 1 cm after a single step of Damped Least Squares optimization and to sub-1 mm after 3 steps in most cases, and can efficiently sample the null space to yield varied configurations for the same pose.
MorphIK is a flow-matching model that encodes a robot's morphology together with the target pose using a transformer and conditions a flow-matching head on that encoding to generate joint poses from noise, thereby solving inverse kinematics for revolute-joint kinematic chains never seen during training; trained only on purely synthetic data from procedurally generated robots, it reaches about 5 cm precision on unseen real-world robots with 6 to 9 Degrees of Freedom, serves as a prior that reduces error to less than 1 cm after a single step of Damped Least Squares optimization and to sub-1 mm after 3 steps in most cases, and can efficiently sample the null space to yield varied configurations for the same pose.
MorphIK is a flow-matching model that encodes a robot's morphology together with the target pose using a transformer and conditions a flow-matching head on that encoding to generate joint poses from noise, thereby solving inverse kinematics for revolute-joint kinematic chains never seen during training; trained only on purely synthetic data from procedurally generated robots, it reaches about 5 cm precision on unseen real-world robots with 6 to 9 Degrees of Freedom, serves as a prior that reduces error to less than 1 cm after a single step of Damped Least Squares optimization and to sub-1 mm after 3 steps in most cases, and can efficiently sample the null space to yield varied configurations for the same pose.
arXiv The work introduces nonorthogonal variational quantum simulation (NOVQS), which applies linear combinations of parameterized quantum states to real- and imaginary-time evolutions, designs a shallow hardware-friendly ansatz tailored to second-quantized electronic-structure Hamiltonians together with resource-efficient protocols for measuring the matrices and vectors in the parameter equations of motion, and provides error analysis and resource estimation; numerical simulations of hydrogen chains and the nitrogen molecule show that a collection of shallow, or even single-layer, parameterized quantum circuits can match or outperform a much deeper circuit in variational quantum simulation, revealing a trade-off between circuit number and depth.
The work introduces nonorthogonal variational quantum simulation (NOVQS), which applies linear combinations of parameterized quantum states to real- and imaginary-time evolutions, designs a shallow hardware-friendly ansatz tailored to second-quantized electronic-structure Hamiltonians together with resource-efficient protocols for measuring the matrices and vectors in the parameter equations of motion, and provides error analysis and resource estimation; numerical simulations of hydrogen chains and the nitrogen molecule show that a collection of shallow, or even single-layer, parameterized quantum circuits can match or outperform a much deeper circuit in variational quantum simulation, revealing a trade-off between circuit number and depth.
The work introduces nonorthogonal variational quantum simulation (NOVQS), which applies linear combinations of parameterized quantum states to real- and imaginary-time evolutions, designs a shallow hardware-friendly ansatz tailored to second-quantized electronic-structure Hamiltonians together with resource-efficient protocols for measuring the matrices and vectors in the parameter equations of motion, and provides error analysis and resource estimation; numerical simulations of hydrogen chains and the nitrogen molecule show that a collection of shallow, or even single-layer, parameterized quantum circuits can match or outperform a much deeper circuit in variational quantum simulation, revealing a trade-off between circuit number and depth.
The work introduces nonorthogonal variational quantum simulation (NOVQS), which applies linear combinations of parameterized quantum states to real- and imaginary-time evolutions, designs a shallow hardware-friendly ansatz tailored to second-quantized electronic-structure Hamiltonians together with resource-efficient protocols for measuring the matrices and vectors in the parameter equations of motion, and provides error analysis and resource estimation; numerical simulations of hydrogen chains and the nitrogen molecule show that a collection of shallow, or even single-layer, parameterized quantum circuits can match or outperform a much deeper circuit in variational quantum simulation, revealing a trade-off between circuit number and depth.
arXiv The work proposes Proof-Sketch-Guided Formal Model Synthesis (ProGS), an autoformalization method centered on model-based proof sketches: a sketch represents the proof structure of the target formal system as a tree, with internal nodes capturing case splits and inductive reasoning steps and leaf nodes corresponding to concrete state-transition events that realize individual subgoals; LLMs generate and repair these sketches, and verification failures are mapped back to specific nodes and subtrees to provide structured guidance for iterative repair; evaluation on a benchmark of 27 formal systems shows ProGS improves over state-of-the-art agentic formal modeling approaches in syntactic validity, deductive verifiability, and behavioral correctness.
The work proposes Proof-Sketch-Guided Formal Model Synthesis (ProGS), an autoformalization method centered on model-based proof sketches: a sketch represents the proof structure of the target formal system as a tree, with internal nodes capturing case splits and inductive reasoning steps and leaf nodes corresponding to concrete state-transition events that realize individual subgoals; LLMs generate and repair these sketches, and verification failures are mapped back to specific nodes and subtrees to provide structured guidance for iterative repair; evaluation on a benchmark of 27 formal systems shows ProGS improves over state-of-the-art agentic formal modeling approaches in syntactic validity, deductive verifiability, and behavioral correctness.
The work proposes Proof-Sketch-Guided Formal Model Synthesis (ProGS), an autoformalization method centered on model-based proof sketches: a sketch represents the proof structure of the target formal system as a tree, with internal nodes capturing case splits and inductive reasoning steps and leaf nodes corresponding to concrete state-transition events that realize individual subgoals; LLMs generate and repair these sketches, and verification failures are mapped back to specific nodes and subtrees to provide structured guidance for iterative repair; evaluation on a benchmark of 27 formal systems shows ProGS improves over state-of-the-art agentic formal modeling approaches in syntactic validity, deductive verifiability, and behavioral correctness.
The work proposes Proof-Sketch-Guided Formal Model Synthesis (ProGS), an autoformalization method centered on model-based proof sketches: a sketch represents the proof structure of the target formal system as a tree, with internal nodes capturing case splits and inductive reasoning steps and leaf nodes corresponding to concrete state-transition events that realize individual subgoals; LLMs generate and repair these sketches, and verification failures are mapped back to specific nodes and subtrees to provide structured guidance for iterative repair; evaluation on a benchmark of 27 formal systems shows ProGS improves over state-of-the-art agentic formal modeling approaches in syntactic validity, deductive verifiability, and behavioral correctness.
arXiv The work introduces Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration that flags reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature, and evaluates it across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test; results reveal a systematic discrepancy between conventional fidelity metrics and CDP, with controlled experiments showing CDP changes monotonically as cross-country divergence is attenuated or amplified while corresponding JSD changes remain relatively small, and an audit of real LLM generations showing DeepPersona-Inspired prompting is frequently favored by conventional fidelity m
The work introduces Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration that flags reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature, and evaluates it across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test; results reveal a systematic discrepancy between conventional fidelity metrics and CDP, with controlled experiments showing CDP changes monotonically as cross-country divergence is attenuated or amplified while corresponding JSD changes remain relatively small, and an audit of real LLM generations showing DeepPersona-Inspired prompting is frequently favored by conventional fidelity m
The work introduces Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration that flags reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature, and evaluates it across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test; results reveal a systematic discrepancy between conventional fidelity metrics and CDP, with controlled experiments showing CDP changes monotonically as cross-country divergence is attenuated or amplified while corresponding JSD changes remain relatively small, and an audit of real LLM generations showing DeepPersona-Inspired prompting is frequently favored by conventional fidelity m
The work introduces Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration that flags reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature, and evaluates it across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test; results reveal a systematic discrepancy between conventional fidelity metrics and CDP, with controlled experiments showing CDP changes monotonically as cross-country divergence is attenuated or amplified while corresponding JSD changes remain relatively small, and an audit of real LLM generations showing DeepPersona-Inspired prompting is frequently favored by conventional fidelity m
arXiv ArGuard is a shared task on harmful content detection in Arabic, with two tracks: Track A on multimodal hate detection in Arabic memes and Track B on harmful prompt detection for Arabic LLM safety evaluation; 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers, with teams exploring models such as AraBERT, Jais, and Qwen3-VL, and the best systems achieving macro-F1 of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2, where fine-grained meme classification in A2 was the most challenging setting due to sparse labels and train-test distribution shifts.
ArGuard is a shared task on harmful content detection in Arabic, with two tracks: Track A on multimodal hate detection in Arabic memes and Track B on harmful prompt detection for Arabic LLM safety evaluation; 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers, with teams exploring models such as AraBERT, Jais, and Qwen3-VL, and the best systems achieving macro-F1 of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2, where fine-grained meme classification in A2 was the most challenging setting due to sparse labels and train-test distribution shifts.
ArGuard is a shared task on harmful content detection in Arabic, with two tracks: Track A on multimodal hate detection in Arabic memes and Track B on harmful prompt detection for Arabic LLM safety evaluation; 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers, with teams exploring models such as AraBERT, Jais, and Qwen3-VL, and the best systems achieving macro-F1 of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2, where fine-grained meme classification in A2 was the most challenging setting due to sparse labels and train-test distribution shifts.
ArGuard is a shared task on harmful content detection in Arabic, with two tracks: Track A on multimodal hate detection in Arabic memes and Track B on harmful prompt detection for Arabic LLM safety evaluation; 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers, with teams exploring models such as AraBERT, Jais, and Qwen3-VL, and the best systems achieving macro-F1 of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2, where fine-grained meme classification in A2 was the most challenging setting due to sparse labels and train-test distribution shifts.
arXiv This review surveys how large language models add semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows, organizing the methods into two operational modes: surrogate models that treat structure-spectrum mapping as a language task, and agentic systems demonstrated to generate code, orchestrate selected simulation steps, and support closed-loop optimization; it also traces the development from classical neural networks to transformer-based models and briefly explores cross-disciplinary applications in fields such as materials science and wireless communications, looking ahead to next-generation multimodal foundation models with physical perception.
This review surveys how large language models add semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows, organizing the methods into two operational modes: surrogate models that treat structure-spectrum mapping as a language task, and agentic systems demonstrated to generate code, orchestrate selected simulation steps, and support closed-loop optimization; it also traces the development from classical neural networks to transformer-based models and briefly explores cross-disciplinary applications in fields such as materials science and wireless communications, looking ahead to next-generation multimodal foundation models with physical perception.
This review surveys how large language models add semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows, organizing the methods into two operational modes: surrogate models that treat structure-spectrum mapping as a language task, and agentic systems demonstrated to generate code, orchestrate selected simulation steps, and support closed-loop optimization; it also traces the development from classical neural networks to transformer-based models and briefly explores cross-disciplinary applications in fields such as materials science and wireless communications, looking ahead to next-generation multimodal foundation models with physical perception.
This review surveys how large language models add semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows, organizing the methods into two operational modes: surrogate models that treat structure-spectrum mapping as a language task, and agentic systems demonstrated to generate code, orchestrate selected simulation steps, and support closed-loop optimization; it also traces the development from classical neural networks to transformer-based models and briefly explores cross-disciplinary applications in fields such as materials science and wireless communications, looking ahead to next-generation multimodal foundation models with physical perception.
arXiv In professional development workshops, 27 middle school teachers used a teacher-facing chatbot authoring tool, and by analyzing focus-group interviews alongside configuration and interaction logs the study examined how pedagogical intentions are translated into chatbot configurations and reflected in chatbot behavior, finding that teachers envisioned chatbots as instructional scaffolds offering differentiated support, extending access to assistance, and preserving student thinking within teacher-defined boundaries; configuration analysis showed Purpose primarily captured instructional goals and content focus while Rules more often specified pedagogical behavior, guardrails, and learner-specific adaptations; and log-based evaluation showed stronger alignment for responsiveness (88.
In professional development workshops, 27 middle school teachers used a teacher-facing chatbot authoring tool, and by analyzing focus-group interviews alongside configuration and interaction logs the study examined how pedagogical intentions are translated into chatbot configurations and reflected in chatbot behavior, finding that teachers envisioned chatbots as instructional scaffolds offering differentiated support, extending access to assistance, and preserving student thinking within teacher-defined boundaries; configuration analysis showed Purpose primarily captured instructional goals and content focus while Rules more often specified pedagogical behavior, guardrails, and learner-specific adaptations; and log-based evaluation showed stronger alignment for responsiveness (88.
In professional development workshops, 27 middle school teachers used a teacher-facing chatbot authoring tool, and by analyzing focus-group interviews alongside configuration and interaction logs the study examined how pedagogical intentions are translated into chatbot configurations and reflected in chatbot behavior, finding that teachers envisioned chatbots as instructional scaffolds offering differentiated support, extending access to assistance, and preserving student thinking within teacher-defined boundaries; configuration analysis showed Purpose primarily captured instructional goals and content focus while Rules more often specified pedagogical behavior, guardrails, and learner-specific adaptations; and log-based evaluation showed stronger alignment for responsiveness (88.
In professional development workshops, 27 middle school teachers used a teacher-facing chatbot authoring tool, and by analyzing focus-group interviews alongside configuration and interaction logs the study examined how pedagogical intentions are translated into chatbot configurations and reflected in chatbot behavior, finding that teachers envisioned chatbots as instructional scaffolds offering differentiated support, extending access to assistance, and preserving student thinking within teacher-defined boundaries; configuration analysis showed Purpose primarily captured instructional goals and content focus while Rules more often specified pedagogical behavior, guardrails, and learner-specific adaptations; and log-based evaluation showed stronger alignment for responsiveness (88.
arXiv The work presents a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers: a shared task-conditioned RL policy handles waypoint-guided approach and loaded transport, an IL policy learns vision-based digging and lifting from expert demonstrations, and digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints; deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control, offline replay and physical experiments show more consistent target selection, shorter local motion time, and increased payload, with the learned digging policy achieving a mean payload of 6.52 kg per completed cycle versus 2.
The work presents a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers: a shared task-conditioned RL policy handles waypoint-guided approach and loaded transport, an IL policy learns vision-based digging and lifting from expert demonstrations, and digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints; deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control, offline replay and physical experiments show more consistent target selection, shorter local motion time, and increased payload, with the learned digging policy achieving a mean payload of 6.52 kg per completed cycle versus 2.
The work presents a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers: a shared task-conditioned RL policy handles waypoint-guided approach and loaded transport, an IL policy learns vision-based digging and lifting from expert demonstrations, and digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints; deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control, offline replay and physical experiments show more consistent target selection, shorter local motion time, and increased payload, with the learned digging policy achieving a mean payload of 6.52 kg per completed cycle versus 2.
The work presents a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers: a shared task-conditioned RL policy handles waypoint-guided approach and loaded transport, an IL policy learns vision-based digging and lifting from expert demonstrations, and digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints; deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control, offline replay and physical experiments show more consistent target selection, shorter local motion time, and increased payload, with the learned digging policy achieving a mean payload of 6.52 kg per completed cycle versus 2.
arXiv The authors propose TabSieve, a select-then-predict framework in which, given a table and a query row, the model first selects a small set of informative rows as evidence and then predicts the missing target conditioned on that evidence; to enable this capability they build TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables with a strong teacher model and strict filtering, and introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards and stabilizes mixed regression and classification training via dynamic task-advantage balancing; on a held-out benchmark of 75 classification and 52 regression tables, TabSieve consistently improves performance across shot budgets, with
The authors propose TabSieve, a select-then-predict framework in which, given a table and a query row, the model first selects a small set of informative rows as evidence and then predicts the missing target conditioned on that evidence; to enable this capability they build TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables with a strong teacher model and strict filtering, and introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards and stabilizes mixed regression and classification training via dynamic task-advantage balancing; on a held-out benchmark of 75 classification and 52 regression tables, TabSieve consistently improves performance across shot budgets, with
The authors propose TabSieve, a select-then-predict framework in which, given a table and a query row, the model first selects a small set of informative rows as evidence and then predicts the missing target conditioned on that evidence; to enable this capability they build TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables with a strong teacher model and strict filtering, and introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards and stabilizes mixed regression and classification training via dynamic task-advantage balancing; on a held-out benchmark of 75 classification and 52 regression tables, TabSieve consistently improves performance across shot budgets, with
The authors propose TabSieve, a select-then-predict framework in which, given a table and a query row, the model first selects a small set of informative rows as evidence and then predicts the missing target conditioned on that evidence; to enable this capability they build TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables with a strong teacher model and strict filtering, and introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards and stabilizes mixed regression and classification training via dynamic task-advantage balancing; on a held-out benchmark of 75 classification and 52 regression tables, TabSieve consistently improves performance across shot budgets, with
arXiv The work introduces CANOPY, a multi-fidelity tree bandit that uses cheap random-path probes to build an online certificate of local aggregation bias and directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation, rather than assuming a global smoothness prior; the authors prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense; across routing, top-k identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including 2.9x higher top-10 recall on a 1000-model pool, 1.
The work introduces CANOPY, a multi-fidelity tree bandit that uses cheap random-path probes to build an online certificate of local aggregation bias and directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation, rather than assuming a global smoothness prior; the authors prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense; across routing, top-k identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including 2.9x higher top-10 recall on a 1000-model pool, 1.
The work introduces CANOPY, a multi-fidelity tree bandit that uses cheap random-path probes to build an online certificate of local aggregation bias and directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation, rather than assuming a global smoothness prior; the authors prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense; across routing, top-k identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including 2.9x higher top-10 recall on a 1000-model pool, 1.
The work introduces CANOPY, a multi-fidelity tree bandit that uses cheap random-path probes to build an online certificate of local aggregation bias and directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation, rather than assuming a global smoothness prior; the authors prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense; across routing, top-k identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including 2.9x higher top-10 recall on a 1000-model pool, 1.
arXiv This work collects roughly 3000 submissions from a couple of Codeforces users, pairs each buggy submission with its corresponding human fix, uses the similarity between the buggy solution and the human fix as a baseline, evaluates the quality of LLM-generated bug fixes on three OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1), and checks whether generated solutions solve the problem using the Codeforces-R1 dataset; the findings suggest LLMs tend to modify more lines than human fixes and sometimes generate entirely new solutions, and that LLMs solve more problems correctly when allowed to generate from scratch rather than patch buggy submissions, even when those submissions are close to the human patch.
This work collects roughly 3000 submissions from a couple of Codeforces users, pairs each buggy submission with its corresponding human fix, uses the similarity between the buggy solution and the human fix as a baseline, evaluates the quality of LLM-generated bug fixes on three OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1), and checks whether generated solutions solve the problem using the Codeforces-R1 dataset; the findings suggest LLMs tend to modify more lines than human fixes and sometimes generate entirely new solutions, and that LLMs solve more problems correctly when allowed to generate from scratch rather than patch buggy submissions, even when those submissions are close to the human patch.
This work collects roughly 3000 submissions from a couple of Codeforces users, pairs each buggy submission with its corresponding human fix, uses the similarity between the buggy solution and the human fix as a baseline, evaluates the quality of LLM-generated bug fixes on three OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1), and checks whether generated solutions solve the problem using the Codeforces-R1 dataset; the findings suggest LLMs tend to modify more lines than human fixes and sometimes generate entirely new solutions, and that LLMs solve more problems correctly when allowed to generate from scratch rather than patch buggy submissions, even when those submissions are close to the human patch.
This work collects roughly 3000 submissions from a couple of Codeforces users, pairs each buggy submission with its corresponding human fix, uses the similarity between the buggy solution and the human fix as a baseline, evaluates the quality of LLM-generated bug fixes on three OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1), and checks whether generated solutions solve the problem using the Codeforces-R1 dataset; the findings suggest LLMs tend to modify more lines than human fixes and sometimes generate entirely new solutions, and that LLMs solve more problems correctly when allowed to generate from scratch rather than patch buggy submissions, even when those submissions are close to the human patch.
arXiv The work proposes STAM-ASR, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR without relying on an external diarization system, learning speaker activity and speaker-aware representations directly from intermediate AudioLLM features to provide explicit who-and-when cues that modulate the semantic representation, while maintaining fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits and that the gap between reference and predicted speaker activity identifies robust speaker tracking
The work proposes STAM-ASR, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR without relying on an external diarization system, learning speaker activity and speaker-aware representations directly from intermediate AudioLLM features to provide explicit who-and-when cues that modulate the semantic representation, while maintaining fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits and that the gap between reference and predicted speaker activity identifies robust speaker tracking
The work proposes STAM-ASR, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR without relying on an external diarization system, learning speaker activity and speaker-aware representations directly from intermediate AudioLLM features to provide explicit who-and-when cues that modulate the semantic representation, while maintaining fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits and that the gap between reference and predicted speaker activity identifies robust speaker tracking
The work proposes STAM-ASR, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR without relying on an external diarization system, learning speaker activity and speaker-aware representations directly from intermediate AudioLLM features to provide explicit who-and-when cues that modulate the semantic representation, while maintaining fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits and that the gap between reference and predicted speaker activity identifies robust speaker tracking
arXiv The work supervises every intermediate depth of one causal speech enhancement model and fine-tunes its output heads so deeper outputs are never worse than shallower ones, yielding a family of static models that are more Pareto-efficient than equivalently-sized counterparts trained from scratch on the same budget (up to 0.11 higher PESQ at equivalent compute, matching the best PESQ at 30% less compute); after int8 quantization on an STM32N6 microcontroller, the dynamic enhancer lies on the same latency-quality frontier as the static models, with the policy running in 26 microseconds per frame on the companion Cortex-M55 and splitting the enhancer into separate NPU graphs adding 2.2% latency overhead.
The work supervises every intermediate depth of one causal speech enhancement model and fine-tunes its output heads so deeper outputs are never worse than shallower ones, yielding a family of static models that are more Pareto-efficient than equivalently-sized counterparts trained from scratch on the same budget (up to 0.11 higher PESQ at equivalent compute, matching the best PESQ at 30% less compute); after int8 quantization on an STM32N6 microcontroller, the dynamic enhancer lies on the same latency-quality frontier as the static models, with the policy running in 26 microseconds per frame on the companion Cortex-M55 and splitting the enhancer into separate NPU graphs adding 2.2% latency overhead.
The work supervises every intermediate depth of one causal speech enhancement model and fine-tunes its output heads so deeper outputs are never worse than shallower ones, yielding a family of static models that are more Pareto-efficient than equivalently-sized counterparts trained from scratch on the same budget (up to 0.11 higher PESQ at equivalent compute, matching the best PESQ at 30% less compute); after int8 quantization on an STM32N6 microcontroller, the dynamic enhancer lies on the same latency-quality frontier as the static models, with the policy running in 26 microseconds per frame on the companion Cortex-M55 and splitting the enhancer into separate NPU graphs adding 2.2% latency overhead.
The work supervises every intermediate depth of one causal speech enhancement model and fine-tunes its output heads so deeper outputs are never worse than shallower ones, yielding a family of static models that are more Pareto-efficient than equivalently-sized counterparts trained from scratch on the same budget (up to 0.11 higher PESQ at equivalent compute, matching the best PESQ at 30% less compute); after int8 quantization on an STM32N6 microcontroller, the dynamic enhancer lies on the same latency-quality frontier as the static models, with the policy running in 26 microseconds per frame on the companion Cortex-M55 and splitting the enhancer into separate NPU graphs adding 2.2% latency overhead.
arXiv The work proposes a training-free method that uses fingertip and toe keypoints from a frozen Sapiens pose foundation model, combined with a per-frame proximity test against annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, to detect which holds a climber uses and when; on the The Way Up dataset (22 videos, 10 athletes, two routes) it reaches an event F1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, performs best on footholds (F1 89.8% overall, 96.
The work proposes a training-free method that uses fingertip and toe keypoints from a frozen Sapiens pose foundation model, combined with a per-frame proximity test against annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, to detect which holds a climber uses and when; on the The Way Up dataset (22 videos, 10 athletes, two routes) it reaches an event F1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, performs best on footholds (F1 89.8% overall, 96.
The work proposes a training-free method that uses fingertip and toe keypoints from a frozen Sapiens pose foundation model, combined with a per-frame proximity test against annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, to detect which holds a climber uses and when; on the The Way Up dataset (22 videos, 10 athletes, two routes) it reaches an event F1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, performs best on footholds (F1 89.8% overall, 96.
The work proposes a training-free method that uses fingertip and toe keypoints from a frozen Sapiens pose foundation model, combined with a per-frame proximity test against annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, to detect which holds a climber uses and when; on the The Way Up dataset (22 videos, 10 athletes, two routes) it reaches an event F1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, performs best on footholds (F1 89.8% overall, 96.
arXiv This study presents a large-scale empirical analysis of cross-OS portability issues across 2,042 open-source Python repositories, combining cross-OS test reexecution (500 projects, 11.2% showing OS-dependent test failures) with manual analysis of 240 GitHub issues (confirming 102 genuine portability problems across 95 additional projects), and develops a taxonomy of 7 primary categories, 24 sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns, while evaluating existing static analysis tools as providing minimal support and large language models achieving 40-79% accuracy in identifying issues and 50-77% success in generating fixes under structured guidance, with practical applicability demonstrated through 33 contributed pull requests (17 merged, zero rejected).
This study presents a large-scale empirical analysis of cross-OS portability issues across 2,042 open-source Python repositories, combining cross-OS test reexecution (500 projects, 11.2% showing OS-dependent test failures) with manual analysis of 240 GitHub issues (confirming 102 genuine portability problems across 95 additional projects), and develops a taxonomy of 7 primary categories, 24 sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns, while evaluating existing static analysis tools as providing minimal support and large language models achieving 40-79% accuracy in identifying issues and 50-77% success in generating fixes under structured guidance, with practical applicability demonstrated through 33 contributed pull requests (17 merged, zero rejected).
This study presents a large-scale empirical analysis of cross-OS portability issues across 2,042 open-source Python repositories, combining cross-OS test reexecution (500 projects, 11.2% showing OS-dependent test failures) with manual analysis of 240 GitHub issues (confirming 102 genuine portability problems across 95 additional projects), and develops a taxonomy of 7 primary categories, 24 sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns, while evaluating existing static analysis tools as providing minimal support and large language models achieving 40-79% accuracy in identifying issues and 50-77% success in generating fixes under structured guidance, with practical applicability demonstrated through 33 contributed pull requests (17 merged, zero rejected).
This study presents a large-scale empirical analysis of cross-OS portability issues across 2,042 open-source Python repositories, combining cross-OS test reexecution (500 projects, 11.2% showing OS-dependent test failures) with manual analysis of 240 GitHub issues (confirming 102 genuine portability problems across 95 additional projects), and develops a taxonomy of 7 primary categories, 24 sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns, while evaluating existing static analysis tools as providing minimal support and large language models achieving 40-79% accuracy in identifying issues and 50-77% success in generating fixes under structured guidance, with practical applicability demonstrated through 33 contributed pull requests (17 merged, zero rejected).
arXiv This paper presents an in-situ qualitative study of a persistent, proactive AI agent 'teammate' deployed across multiple teams in a large technology company, finding that the boundaries of the human-agent workplace are actively in flux and trigger breakdowns and negotiations across three areas: tacit rules of collaborative human workflows, the relational boundaries of this new non-human actor, and the redistribution of trust and human agency; the authors use these early micro-negotiations as signals to chart a research, design, and organizational agenda that intentionally preserves human agency.
This paper presents an in-situ qualitative study of a persistent, proactive AI agent 'teammate' deployed across multiple teams in a large technology company, finding that the boundaries of the human-agent workplace are actively in flux and trigger breakdowns and negotiations across three areas: tacit rules of collaborative human workflows, the relational boundaries of this new non-human actor, and the redistribution of trust and human agency; the authors use these early micro-negotiations as signals to chart a research, design, and organizational agenda that intentionally preserves human agency.
This paper presents an in-situ qualitative study of a persistent, proactive AI agent 'teammate' deployed across multiple teams in a large technology company, finding that the boundaries of the human-agent workplace are actively in flux and trigger breakdowns and negotiations across three areas: tacit rules of collaborative human workflows, the relational boundaries of this new non-human actor, and the redistribution of trust and human agency; the authors use these early micro-negotiations as signals to chart a research, design, and organizational agenda that intentionally preserves human agency.
This paper presents an in-situ qualitative study of a persistent, proactive AI agent 'teammate' deployed across multiple teams in a large technology company, finding that the boundaries of the human-agent workplace are actively in flux and trigger breakdowns and negotiations across three areas: tacit rules of collaborative human workflows, the relational boundaries of this new non-human actor, and the redistribution of trust and human agency; the authors use these early micro-negotiations as signals to chart a research, design, and organizational agenda that intentionally preserves human agency.
arXiv The study hypothesizes that agent authorization can be formalized as a cryptographically verifiable relation, R_CVA, that jointly binds an agent principal, a concrete authorization request, an execution context, and the satisfaction of an applicable policy while selectively preserving the confidentiality of private authorization attributes; it introduces a preliminary formal abstraction for Cryptographically Verifiable Agent Authorization (CVA), defines a compact set of candidate security properties including authorization soundness, principal binding, request binding, policy binding, and replay resistance, provides an executable zero-knowledge proof of concept instantiating selected elements of the model over a Groth16 zk-SNARK construction, and formalizes the structural separation among
The study hypothesizes that agent authorization can be formalized as a cryptographically verifiable relation, R_CVA, that jointly binds an agent principal, a concrete authorization request, an execution context, and the satisfaction of an applicable policy while selectively preserving the confidentiality of private authorization attributes; it introduces a preliminary formal abstraction for Cryptographically Verifiable Agent Authorization (CVA), defines a compact set of candidate security properties including authorization soundness, principal binding, request binding, policy binding, and replay resistance, provides an executable zero-knowledge proof of concept instantiating selected elements of the model over a Groth16 zk-SNARK construction, and formalizes the structural separation among
The study hypothesizes that agent authorization can be formalized as a cryptographically verifiable relation, R_CVA, that jointly binds an agent principal, a concrete authorization request, an execution context, and the satisfaction of an applicable policy while selectively preserving the confidentiality of private authorization attributes; it introduces a preliminary formal abstraction for Cryptographically Verifiable Agent Authorization (CVA), defines a compact set of candidate security properties including authorization soundness, principal binding, request binding, policy binding, and replay resistance, provides an executable zero-knowledge proof of concept instantiating selected elements of the model over a Groth16 zk-SNARK construction, and formalizes the structural separation among
The study hypothesizes that agent authorization can be formalized as a cryptographically verifiable relation, R_CVA, that jointly binds an agent principal, a concrete authorization request, an execution context, and the satisfaction of an applicable policy while selectively preserving the confidentiality of private authorization attributes; it introduces a preliminary formal abstraction for Cryptographically Verifiable Agent Authorization (CVA), defines a compact set of candidate security properties including authorization soundness, principal binding, request binding, policy binding, and replay resistance, provides an executable zero-knowledge proof of concept instantiating selected elements of the model over a Groth16 zk-SNARK construction, and formalizes the structural separation among
arXiv The authors introduce MemGuard-Alpha, comprising a composite contamination score combining five membership inference attack (MIA) methods with a temporal proximity feature, plus Cross-Model Memorization Disagreement exploiting variation in training cutoffs across models, and audit both across seven LLMs (124M–7B), 50 S&P 100 constituents, 42,800 prompts and 299,600 prompt-model MIA scores spanning 2019–2024, yielding three negative results: the temporal proximity feature recovers the in-sample label perfectly (ROC-AUC 1.000) because it is a monotone transform of the defining variable, the discriminative power of MIA scores is largely attributable to scale differences between models (raw scores reach AUC up to 0.99 at a fixed date but fall to 0.487–0.
The authors introduce MemGuard-Alpha, comprising a composite contamination score combining five membership inference attack (MIA) methods with a temporal proximity feature, plus Cross-Model Memorization Disagreement exploiting variation in training cutoffs across models, and audit both across seven LLMs (124M–7B), 50 S&P 100 constituents, 42,800 prompts and 299,600 prompt-model MIA scores spanning 2019–2024, yielding three negative results: the temporal proximity feature recovers the in-sample label perfectly (ROC-AUC 1.000) because it is a monotone transform of the defining variable, the discriminative power of MIA scores is largely attributable to scale differences between models (raw scores reach AUC up to 0.99 at a fixed date but fall to 0.487–0.
The authors introduce MemGuard-Alpha, comprising a composite contamination score combining five membership inference attack (MIA) methods with a temporal proximity feature, plus Cross-Model Memorization Disagreement exploiting variation in training cutoffs across models, and audit both across seven LLMs (124M–7B), 50 S&P 100 constituents, 42,800 prompts and 299,600 prompt-model MIA scores spanning 2019–2024, yielding three negative results: the temporal proximity feature recovers the in-sample label perfectly (ROC-AUC 1.000) because it is a monotone transform of the defining variable, the discriminative power of MIA scores is largely attributable to scale differences between models (raw scores reach AUC up to 0.99 at a fixed date but fall to 0.487–0.
The authors introduce MemGuard-Alpha, comprising a composite contamination score combining five membership inference attack (MIA) methods with a temporal proximity feature, plus Cross-Model Memorization Disagreement exploiting variation in training cutoffs across models, and audit both across seven LLMs (124M–7B), 50 S&P 100 constituents, 42,800 prompts and 299,600 prompt-model MIA scores spanning 2019–2024, yielding three negative results: the temporal proximity feature recovers the in-sample label perfectly (ROC-AUC 1.000) because it is a monotone transform of the defining variable, the discriminative power of MIA scores is largely attributable to scale differences between models (raw scores reach AUC up to 0.99 at a fixed date but fall to 0.487–0.