Skip to main content

Search

“All disciplines” · 1495 results

Page 34 · showing 20
arXiv

DMA approximates factor-to-variable messages directly, bounds marginal KL by message KL, and yields a Bayesian neural network inference algorithm without a learning-rate hyperparameter

The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
arXiv

SciMuse generates personalized research ideas from a 58-million-paper knowledge graph and an LLM; over 100 research group leaders rated 4,400+ ideas at a mean of 2.40 on a 5-point scale, with 24.9% rated 4 or 5

The authors introduce SciMuse, which generates personalized research ideas using a knowledge graph of 58 million papers and a large language model, and they conducted a large-scale evaluation in which more than 100 research group leaders spanning the natural sciences to the humanities rated over 4,400 personalized ideas by level of interest; expert ratings were modest (mean 2.40 on a 5-point scale, most common rating 1) while 24.9% of ideas were rated 4 or 5, supplying knowledge-graph-selected concept pairs did not improve expert-rated interest over a titles-only GPT baseline, high-citation-predicted pairs even showed a weak tendency (1.
arXiv

MorphIK solves inverse kinematics for unseen 6-to-9-DoF robots with morphology-conditioned flow matching, reaching about 5 cm and sub-millimeter accuracy after optimization

MorphIK is a flow-matching model that encodes a robot's morphology together with the target pose using a transformer and conditions a flow-matching head on that encoding to generate joint poses from noise, thereby solving inverse kinematics for revolute-joint kinematic chains never seen during training; trained only on purely synthetic data from procedurally generated robots, it reaches about 5 cm precision on unseen real-world robots with 6 to 9 Degrees of Freedom, serves as a prior that reduces error to less than 1 cm after a single step of Damped Least Squares optimization and to sub-1 mm after 3 steps in most cases, and can efficiently sample the null space to yield varied configurations for the same pose.
arXiv

NOVQS combines multiple shallow parameterized quantum circuits and matches or outperforms a single deeper circuit in hydrogen-chain and nitrogen-molecule simulations

The work introduces nonorthogonal variational quantum simulation (NOVQS), which applies linear combinations of parameterized quantum states to real- and imaginary-time evolutions, designs a shallow hardware-friendly ansatz tailored to second-quantized electronic-structure Hamiltonians together with resource-efficient protocols for measuring the matrices and vectors in the parameter equations of motion, and provides error analysis and resource estimation; numerical simulations of hydrogen chains and the nitrogen molecule show that a collection of shallow, or even single-layer, parameterized quantum circuits can match or outperform a much deeper circuit in variational quantum simulation, revealing a trade-off between circuit number and depth.
arXiv

ProGS organizes formal modeling around model-based proof sketches, improving syntactic validity, deductive verifiability, and behavioral correctness on a 27-system benchmark

The work proposes Proof-Sketch-Guided Formal Model Synthesis (ProGS), an autoformalization method centered on model-based proof sketches: a sketch represents the proof structure of the target formal system as a tree, with internal nodes capturing case splits and inductive reasoning steps and leaf nodes corresponding to concrete state-transition events that realize individual subgoals; LLMs generate and repair these sketches, and verification failures are mapped back to specific nodes and subtrees to provide structured guidance for iterative repair; evaluation on a benchmark of 27 formal systems shows ProGS improves over state-of-the-art agentic formal modeling approaches in syntactic validity, deductive verifiability, and behavioral correctness.
arXiv

A new CDP diagnostic finds DeepPersona-Inspired prompting shows the strongest cultural flattening in every model-domain block across four backbones and two survey domains

The work introduces Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration that flags reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature, and evaluates it across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test; results reveal a systematic discrepancy between conventional fidelity metrics and CDP, with controlled experiments showing CDP changes monotonically as cross-country divergence is attenuated or amplified while corresponding JSD changes remain relatively small, and an audit of real LLM generations showing DeepPersona-Inspired prompting is frequently favored by conventional fidelity m
arXiv

ArGuard Shared Task: 58 teams registered and 35 competed on Arabic meme and LLM-prompt harm detection, with best systems reaching macro-F1 of 0.823/0.419/0.984/0.790

ArGuard is a shared task on harmful content detection in Arabic, with two tracks: Track A on multimodal hate detection in Arabic memes and Track B on harmful prompt detection for Arabic LLM safety evaluation; 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers, with teams exploring models such as AraBERT, Jais, and Qwen3-VL, and the best systems achieving macro-F1 of 0.823 on A1, 0.419 on A2, 0.984 on B1, and 0.790 on B2, where fine-grained meme classification in A2 was the most challenging setting due to sparse labels and train-test distribution shifts.
arXiv

Review maps how large language models enter nanophotonic design as surrogate models and agentic systems

This review surveys how large language models add semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows, organizing the methods into two operational modes: surrogate models that treat structure-spectrum mapping as a language task, and agentic systems demonstrated to generate code, orchestrate selected simulation steps, and support closed-loop optimization; it also traces the development from classical neural networks to transformer-based models and briefly explores cross-disciplinary applications in fields such as materials science and wireless communications, looking ahead to next-generation multimodal foundation models with physical perception.
arXiv

Twenty-seven middle school teachers configured pedagogical intent in a teacher-facing chatbot authoring tool, and log-based evaluation showed 88.9% alignment for responsiveness and 81.5% for persona versus 70.4% for rules and 59.3% for purpose

In professional development workshops, 27 middle school teachers used a teacher-facing chatbot authoring tool, and by analyzing focus-group interviews alongside configuration and interaction logs the study examined how pedagogical intentions are translated into chatbot configurations and reflected in chatbot behavior, finding that teachers envisioned chatbots as instructional scaffolds offering differentiated support, extending access to assistance, and preserving student thinking within teacher-defined boundaries; configuration analysis showed Purpose primarily captured instructional goals and content focus while Rules more often specified pedagogical behavior, guardrails, and learner-specific adaptations; and log-based evaluation showed stronger alignment for responsiveness (88.
arXiv

Learning-based framework enables continuous autonomous excavation on a scaled hydraulic excavator, averaging 6.52 kg payload per cycle versus 2.68 kg for Fixed Dig

The work presents a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers: a shared task-conditioned RL policy handles waypoint-guided approach and loaded transport, an IL policy learns vision-based digging and lifting from expert demonstrations, and digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints; deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control, offline replay and physical experiments show more consistent target selection, shorter local motion time, and increased payload, with the learned digging policy achieving a mean payload of 6.52 kg per completed cycle versus 2.
arXiv

TabSieve's select-then-predict framework lifts classification by 2.92% and regression by 4.45% on average across 75 classification and 52 regression tables

The authors propose TabSieve, a select-then-predict framework in which, given a table and a query row, the model first selects a small set of informative rows as evidence and then predicts the missing target conditioned on that evidence; to enable this capability they build TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables with a strong teacher model and strict filtering, and introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards and stabilizes mixed regression and classification training via dynamic task-advantage balancing; on a held-out benchmark of 75 classification and 52 regression tables, TabSieve consistently improves performance across shot budgets, with
arXiv

CANOPY uses random-path probes to certify smoothness violations online, improving routing, top-k identification, test-time search, caching, and prompt trimming at matched budgets

The work introduces CANOPY, a multi-fidelity tree bandit that uses cheap random-path probes to build an online certificate of local aggregation bias and directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation, rather than assuming a global smoothness prior; the authors prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense; across routing, top-k identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including 2.9x higher top-10 recall on a 1000-model pool, 1.
arXiv

GPT models change more lines than human patches when fixing Codeforces submissions, and solve more problems from scratch than by patching

This work collects roughly 3000 submissions from a couple of Codeforces users, pairs each buggy submission with its corresponding human fix, uses the similarity between the buggy solution and the human fix as a baseline, evaluates the quality of LLM-generated bug fixes on three OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1), and checks whether generated solutions solve the problem using the Codeforces-R1 dataset; the findings suggest LLMs tend to modify more lines than human fixes and sometimes generate entirely new solutions, and that LLMs solve more problems correctly when allowed to generate from scratch rather than patch buggy submissions, even when those submissions are close to the human patch.
arXiv

STAM-ASR extends a pretrained AudioLLM with speaker-temporal anchoring and memory, evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 for multi-speaker ASR

The work proposes STAM-ASR, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR without relying on an external diarization system, learning speaker activity and speaker-aware representations directly from intermediate AudioLLM features to provide explicit who-and-when cues that modulate the semantic representation, while maintaining fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits and that the gap between reference and predicted speaker activity identifies robust speaker tracking
arXiv

Does per-frame early exit pay off? A dynamic-depth speech enhancer lands on the same latency-quality frontier as static models on STM32N6, with a 26 microsecond per-frame policy

The work supervises every intermediate depth of one causal speech enhancement model and fine-tunes its output heads so deeper outputs are never worse than shallower ones, yielding a family of static models that are more Pareto-efficient than equivalently-sized counterparts trained from scratch on the same budget (up to 0.11 higher PESQ at equivalent compute, matching the best PESQ at 30% less compute); after int8 quantization on an STM32N6 microcontroller, the dynamic enhancer lies on the same latency-quality frontier as the static models, with the policy running in 26 microseconds per frame on the companion Cortex-M55 and splitting the enhancer into separate NPU graphs adding 2.2% latency overhead.
arXiv

A frozen Sapiens pose foundation model detects climbing hold usage without climbing-specific training, reaching 90.2% event F1

The work proposes a training-free method that uses fingertip and toe keypoints from a frozen Sapiens pose foundation model, combined with a per-frame proximity test against annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, to detect which holds a climber uses and when; on the The Way Up dataset (22 videos, 10 athletes, two routes) it reaches an event F1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, performs best on footholds (F1 89.8% overall, 96.
arXiv

An empirical analysis of 2,042 Python repositories finds 11.2% show OS-dependent test failures, yielding a 7-category taxonomy and an LLM repair evaluation

This study presents a large-scale empirical analysis of cross-OS portability issues across 2,042 open-source Python repositories, combining cross-OS test reexecution (500 projects, 11.2% showing OS-dependent test failures) with manual analysis of 240 GitHub issues (confirming 102 genuine portability problems across 95 additional projects), and develops a taxonomy of 7 primary categories, 24 sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns, while evaluating existing static analysis tools as providing minimal support and large language models achieving 40-79% accuracy in identifying issues and 50-77% success in generating fixes under structured guidance, with practical applicability demonstrated through 33 contributed pull requests (17 merged, zero rejected).
arXiv

After deploying a persistent, proactive AI 'teammate' across multiple teams at a large technology company, researchers observed human-agent workplace boundaries being continually negotiated across collaboration rules, the non-human actor's relational boundaries, and the redistribution of trust.

This paper presents an in-situ qualitative study of a persistent, proactive AI agent 'teammate' deployed across multiple teams in a large technology company, finding that the boundaries of the human-agent workplace are actively in flux and trigger breakdowns and negotiations across three areas: tacit rules of collaborative human workflows, the relational boundaries of this new non-human actor, and the redistribution of trust and human agency; the authors use these early micro-negotiations as signals to chart a research, design, and organizational agenda that intentionally preserves human agency.
arXiv

Agent authorization recast as a cryptographically verifiable relation R_CVA, with an executable Groth16 zk-SNARK proof of concept

The study hypothesizes that agent authorization can be formalized as a cryptographically verifiable relation, R_CVA, that jointly binds an agent principal, a concrete authorization request, an execution context, and the satisfaction of an applicable policy while selectively preserving the confidentiality of private authorization attributes; it introduces a preliminary formal abstraction for Cryptographically Verifiable Agent Authorization (CVA), defines a compact set of candidate security properties including authorization soundness, principal binding, request binding, policy binding, and replay resistance, provides an executable zero-knowledge proof of concept instantiating selected elements of the model over a Groth16 zk-SNARK construction, and formalizes the structural separation among
arXiv

MemGuard-Alpha audit finds MIA contamination scores drop to 0.487–0.537 discrimination at a fixed date, and no filtering variant beats the unfiltered ensemble under symmetric transaction costs

The authors introduce MemGuard-Alpha, comprising a composite contamination score combining five membership inference attack (MIA) methods with a temporal proximity feature, plus Cross-Model Memorization Disagreement exploiting variation in training cutoffs across models, and audit both across seven LLMs (124M–7B), 50 S&P 100 constituents, 42,800 prompts and 299,600 prompt-model MIA scores spanning 2019–2024, yielding three negative results: the temporal proximity feature recovers the in-sample label perfectly (ROC-AUC 1.000) because it is a monotone transform of the defining variable, the discriminative power of MIA scores is largely attributable to scale differences between models (raw scores reach AUC up to 0.99 at a fixed date but fall to 0.487–0.