arXiv The work introduces Symbolic Closure Analysis (SCA) to characterize exploration bias and compounding bias in long-horizon reasoning, and builds SAGE, a framework that injects structural priors during post-training via algebraic sparsification and hyperbolic structural guidance, outperforming SFT, GRPO, EMPO, and GRPO-PRM across 12 benchmarks and 7 model families and achieving up to an 8-fold improvement in Lean-verified proofs on the open Andrews-Curtis task.
The work introduces Symbolic Closure Analysis (SCA) to characterize exploration bias and compounding bias in long-horizon reasoning, and builds SAGE, a framework that injects structural priors during post-training via algebraic sparsification and hyperbolic structural guidance, outperforming SFT, GRPO, EMPO, and GRPO-PRM across 12 benchmarks and 7 model families and achieving up to an 8-fold improvement in Lean-verified proofs on the open Andrews-Curtis task.
The work introduces Symbolic Closure Analysis (SCA) to characterize exploration bias and compounding bias in long-horizon reasoning, and builds SAGE, a framework that injects structural priors during post-training via algebraic sparsification and hyperbolic structural guidance, outperforming SFT, GRPO, EMPO, and GRPO-PRM across 12 benchmarks and 7 model families and achieving up to an 8-fold improvement in Lean-verified proofs on the open Andrews-Curtis task.
The work introduces Symbolic Closure Analysis (SCA) to characterize exploration bias and compounding bias in long-horizon reasoning, and builds SAGE, a framework that injects structural priors during post-training via algebraic sparsification and hyperbolic structural guidance, outperforming SFT, GRPO, EMPO, and GRPO-PRM across 12 benchmarks and 7 model families and achieving up to an 8-fold improvement in Lean-verified proofs on the open Andrews-Curtis task.
arXiv ZooWork presents ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, 8B) trained on preference labels produced by three reasoning LLM families (Qwen3.5-122B, Gemma-4-31B, DeepSeek-V4-Pro) under a constraint-first, both-orders judging protocol, with the aligned 8B distilling into the 4B and 0.6B; on ShopRank-Bench, built from Gensmo private traffic (10,511 pairs tiered gold/silver/bronze), the 8B and 4B significantly outperform the strongest open baseline Jina-m0, every size significantly beats its own un-aligned base, the 0.6B is statistically indistinguishable from the 4B base on structured text, and alignment costs nothing in general MTEB reranking quality or serving latency.
ZooWork presents ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, 8B) trained on preference labels produced by three reasoning LLM families (Qwen3.5-122B, Gemma-4-31B, DeepSeek-V4-Pro) under a constraint-first, both-orders judging protocol, with the aligned 8B distilling into the 4B and 0.6B; on ShopRank-Bench, built from Gensmo private traffic (10,511 pairs tiered gold/silver/bronze), the 8B and 4B significantly outperform the strongest open baseline Jina-m0, every size significantly beats its own un-aligned base, the 0.6B is statistically indistinguishable from the 4B base on structured text, and alignment costs nothing in general MTEB reranking quality or serving latency.
ZooWork presents ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, 8B) trained on preference labels produced by three reasoning LLM families (Qwen3.5-122B, Gemma-4-31B, DeepSeek-V4-Pro) under a constraint-first, both-orders judging protocol, with the aligned 8B distilling into the 4B and 0.6B; on ShopRank-Bench, built from Gensmo private traffic (10,511 pairs tiered gold/silver/bronze), the 8B and 4B significantly outperform the strongest open baseline Jina-m0, every size significantly beats its own un-aligned base, the 0.6B is statistically indistinguishable from the 4B base on structured text, and alignment costs nothing in general MTEB reranking quality or serving latency.
ZooWork presents ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, 8B) trained on preference labels produced by three reasoning LLM families (Qwen3.5-122B, Gemma-4-31B, DeepSeek-V4-Pro) under a constraint-first, both-orders judging protocol, with the aligned 8B distilling into the 4B and 0.6B; on ShopRank-Bench, built from Gensmo private traffic (10,511 pairs tiered gold/silver/bronze), the 8B and 4B significantly outperform the strongest open baseline Jina-m0, every size significantly beats its own un-aligned base, the 0.6B is statistically indistinguishable from the 4B base on structured text, and alignment costs nothing in general MTEB reranking quality or serving latency.
arXiv InternW0-Δ couples a pretrained video expert and an action expert through a directed Mixture-of-Transformers, uses a frozen VLM for scene semantics, learns future-relevant scene changes via Causal Imprint from training-only future supervision, and distills 4D geometric and motion priors from a Track4World teacher at training time only; pretrained on a corpus of over 20K hours unifying robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data under a canonical state-action representation, it reaches 92.8% on LIBERO-Plus, 71.9% on RoboTwin 2.0 Clean2Random, an overall score of 66.0 on EBench, and an average success rate of 23.91% on RoboDojo, with deployment on four real-robot platforms (two gripper-based, two dexterous-hand).
InternW0-Δ couples a pretrained video expert and an action expert through a directed Mixture-of-Transformers, uses a frozen VLM for scene semantics, learns future-relevant scene changes via Causal Imprint from training-only future supervision, and distills 4D geometric and motion priors from a Track4World teacher at training time only; pretrained on a corpus of over 20K hours unifying robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data under a canonical state-action representation, it reaches 92.8% on LIBERO-Plus, 71.9% on RoboTwin 2.0 Clean2Random, an overall score of 66.0 on EBench, and an average success rate of 23.91% on RoboDojo, with deployment on four real-robot platforms (two gripper-based, two dexterous-hand).
InternW0-Δ couples a pretrained video expert and an action expert through a directed Mixture-of-Transformers, uses a frozen VLM for scene semantics, learns future-relevant scene changes via Causal Imprint from training-only future supervision, and distills 4D geometric and motion priors from a Track4World teacher at training time only; pretrained on a corpus of over 20K hours unifying robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data under a canonical state-action representation, it reaches 92.8% on LIBERO-Plus, 71.9% on RoboTwin 2.0 Clean2Random, an overall score of 66.0 on EBench, and an average success rate of 23.91% on RoboDojo, with deployment on four real-robot platforms (two gripper-based, two dexterous-hand).
InternW0-Δ couples a pretrained video expert and an action expert through a directed Mixture-of-Transformers, uses a frozen VLM for scene semantics, learns future-relevant scene changes via Causal Imprint from training-only future supervision, and distills 4D geometric and motion priors from a Track4World teacher at training time only; pretrained on a corpus of over 20K hours unifying robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data under a canonical state-action representation, it reaches 92.8% on LIBERO-Plus, 71.9% on RoboTwin 2.0 Clean2Random, an overall score of 66.0 on EBench, and an average success rate of 23.91% on RoboDojo, with deployment on four real-robot platforms (two gripper-based, two dexterous-hand).
arXiv The study presents a large-scale, data-driven analysis of 2,170 publicly available Jev projects collected from GitHub as of September 22, 2026, finding rapid early growth with 1,865 new repositories and 305 integrations into existing repositories within a week of release, projects using Jev for multiple decision purposes and combining its interfaces across domains, with attribute judgment (77%) and scoring or ranking (52%) most common, while public attention concentrates in routing and interface agents (19.6% of projects but 63.0% of stars) and does not track project counts.
The study presents a large-scale, data-driven analysis of 2,170 publicly available Jev projects collected from GitHub as of September 22, 2026, finding rapid early growth with 1,865 new repositories and 305 integrations into existing repositories within a week of release, projects using Jev for multiple decision purposes and combining its interfaces across domains, with attribute judgment (77%) and scoring or ranking (52%) most common, while public attention concentrates in routing and interface agents (19.6% of projects but 63.0% of stars) and does not track project counts.
The study presents a large-scale, data-driven analysis of 2,170 publicly available Jev projects collected from GitHub as of September 22, 2026, finding rapid early growth with 1,865 new repositories and 305 integrations into existing repositories within a week of release, projects using Jev for multiple decision purposes and combining its interfaces across domains, with attribute judgment (77%) and scoring or ranking (52%) most common, while public attention concentrates in routing and interface agents (19.6% of projects but 63.0% of stars) and does not track project counts.
The study presents a large-scale, data-driven analysis of 2,170 publicly available Jev projects collected from GitHub as of September 22, 2026, finding rapid early growth with 1,865 new repositories and 305 integrations into existing repositories within a week of release, projects using Jev for multiple decision purposes and combining its interfaces across domains, with attribute judgment (77%) and scoring or ranking (52%) most common, while public attention concentrates in routing and interface agents (19.6% of projects but 63.0% of stars) and does not track project counts.
arXiv The authors introduce ARGUS, a bounded retrieval-gated large-language-model pipeline that decomposes a difference-in-differences (DID) study into eleven assumption-implication-evidence dimensions, audits whether the evidence a paper reports adequately supports each identification assumption, and abstains when no relevant evidence can be retrieved; it detects 8 of 11 planted flaws (versus 2 for a keyword baseline), 25 of 33 flaw variants, abstains on roughly 60% of paper-dimension assessments over 26 economics papers for lack of retrievable evidence, and in a five-paper, 55-cell pilot with two reconciled annotators is more severe than the human labels on 25 of the 33 cells it completes, an over-severity a rule fixed before the labels arrived reduces substantially in-sample.
The authors introduce ARGUS, a bounded retrieval-gated large-language-model pipeline that decomposes a difference-in-differences (DID) study into eleven assumption-implication-evidence dimensions, audits whether the evidence a paper reports adequately supports each identification assumption, and abstains when no relevant evidence can be retrieved; it detects 8 of 11 planted flaws (versus 2 for a keyword baseline), 25 of 33 flaw variants, abstains on roughly 60% of paper-dimension assessments over 26 economics papers for lack of retrievable evidence, and in a five-paper, 55-cell pilot with two reconciled annotators is more severe than the human labels on 25 of the 33 cells it completes, an over-severity a rule fixed before the labels arrived reduces substantially in-sample.
The authors introduce ARGUS, a bounded retrieval-gated large-language-model pipeline that decomposes a difference-in-differences (DID) study into eleven assumption-implication-evidence dimensions, audits whether the evidence a paper reports adequately supports each identification assumption, and abstains when no relevant evidence can be retrieved; it detects 8 of 11 planted flaws (versus 2 for a keyword baseline), 25 of 33 flaw variants, abstains on roughly 60% of paper-dimension assessments over 26 economics papers for lack of retrievable evidence, and in a five-paper, 55-cell pilot with two reconciled annotators is more severe than the human labels on 25 of the 33 cells it completes, an over-severity a rule fixed before the labels arrived reduces substantially in-sample.
The authors introduce ARGUS, a bounded retrieval-gated large-language-model pipeline that decomposes a difference-in-differences (DID) study into eleven assumption-implication-evidence dimensions, audits whether the evidence a paper reports adequately supports each identification assumption, and abstains when no relevant evidence can be retrieved; it detects 8 of 11 planted flaws (versus 2 for a keyword baseline), 25 of 33 flaw variants, abstains on roughly 60% of paper-dimension assessments over 26 economics papers for lack of retrievable evidence, and in a five-paper, 55-cell pilot with two reconciled annotators is more severe than the human labels on 25 of the 33 cells it completes, an over-severity a rule fixed before the labels arrived reduces substantially in-sample.
arXiv The work introduces and first implements end-to-end continuous depth batching (CDB) for looped language models: it splits inference into prefill, prelude, recurrent-core, and coda stage queues, reforms batches between loop steps, manages a depth-indexed KV cache, and uses a lookahead gate to predict exits one step ahead so tokens with different loop depths can share a forward pass; on Ouro 1.4B and Huginn 3.5B, CDB realizes about 99% and 83%–96% of the estimated maximum speedup, showing fully looped architectures are best suited to depth-adaptive inference.
The work introduces and first implements end-to-end continuous depth batching (CDB) for looped language models: it splits inference into prefill, prelude, recurrent-core, and coda stage queues, reforms batches between loop steps, manages a depth-indexed KV cache, and uses a lookahead gate to predict exits one step ahead so tokens with different loop depths can share a forward pass; on Ouro 1.4B and Huginn 3.5B, CDB realizes about 99% and 83%–96% of the estimated maximum speedup, showing fully looped architectures are best suited to depth-adaptive inference.
The work introduces and first implements end-to-end continuous depth batching (CDB) for looped language models: it splits inference into prefill, prelude, recurrent-core, and coda stage queues, reforms batches between loop steps, manages a depth-indexed KV cache, and uses a lookahead gate to predict exits one step ahead so tokens with different loop depths can share a forward pass; on Ouro 1.4B and Huginn 3.5B, CDB realizes about 99% and 83%–96% of the estimated maximum speedup, showing fully looped architectures are best suited to depth-adaptive inference.
The work introduces and first implements end-to-end continuous depth batching (CDB) for looped language models: it splits inference into prefill, prelude, recurrent-core, and coda stage queues, reforms batches between loop steps, manages a depth-indexed KV cache, and uses a lookahead gate to predict exits one step ahead so tokens with different loop depths can share a forward pass; on Ouro 1.4B and Huginn 3.5B, CDB realizes about 99% and 83%–96% of the estimated maximum speedup, showing fully looped architectures are best suited to depth-adaptive inference.
arXiv The work splits tool-calling RL trajectories into a tool segment and a summary segment, argues that broadcasting a single trajectory-level advantage causes cross-segment credit misattribution, and proposes SLCA-GRPO, which normalizes the two segment rewards within the group and routes each only to its own tokens inside a single unified policy, paired with a Schema-Guided LLM Simulator and Hierarchical Rewards; across Qwen2.5-3B/7B-Instruct and Qwen3-8B-Base on Toucan-Test, BFCL V3 and τ²-Bench it reports higher means than matched GRPO, with 7B gaps of +2.53, +1.36 and +9.15 percentage points.
The work splits tool-calling RL trajectories into a tool segment and a summary segment, argues that broadcasting a single trajectory-level advantage causes cross-segment credit misattribution, and proposes SLCA-GRPO, which normalizes the two segment rewards within the group and routes each only to its own tokens inside a single unified policy, paired with a Schema-Guided LLM Simulator and Hierarchical Rewards; across Qwen2.5-3B/7B-Instruct and Qwen3-8B-Base on Toucan-Test, BFCL V3 and τ²-Bench it reports higher means than matched GRPO, with 7B gaps of +2.53, +1.36 and +9.15 percentage points.
The work splits tool-calling RL trajectories into a tool segment and a summary segment, argues that broadcasting a single trajectory-level advantage causes cross-segment credit misattribution, and proposes SLCA-GRPO, which normalizes the two segment rewards within the group and routes each only to its own tokens inside a single unified policy, paired with a Schema-Guided LLM Simulator and Hierarchical Rewards; across Qwen2.5-3B/7B-Instruct and Qwen3-8B-Base on Toucan-Test, BFCL V3 and τ²-Bench it reports higher means than matched GRPO, with 7B gaps of +2.53, +1.36 and +9.15 percentage points.
The work splits tool-calling RL trajectories into a tool segment and a summary segment, argues that broadcasting a single trajectory-level advantage causes cross-segment credit misattribution, and proposes SLCA-GRPO, which normalizes the two segment rewards within the group and routes each only to its own tokens inside a single unified policy, paired with a Schema-Guided LLM Simulator and Hierarchical Rewards; across Qwen2.5-3B/7B-Instruct and Qwen3-8B-Base on Toucan-Test, BFCL V3 and τ²-Bench it reports higher means than matched GRPO, with 7B gaps of +2.53, +1.36 and +9.15 percentage points.
arXiv The work proposes FoMo: forward-noising a reference image to a sampled timestep and then denoising it independently, using that forking moment as a pointwise perceptual-distance label, so training data can be generated with no human annotation, and trains reference-based IQA metrics with a RankNet-style global ranking objective; across four benchmarks and seven backbones it achieves the best average performance, with LPIPS-Alex reaching 0.733 SROCC on PIPAL versus 0.577 for the same architecture trained on human-annotated KADID-10k labels.
The work proposes FoMo: forward-noising a reference image to a sampled timestep and then denoising it independently, using that forking moment as a pointwise perceptual-distance label, so training data can be generated with no human annotation, and trains reference-based IQA metrics with a RankNet-style global ranking objective; across four benchmarks and seven backbones it achieves the best average performance, with LPIPS-Alex reaching 0.733 SROCC on PIPAL versus 0.577 for the same architecture trained on human-annotated KADID-10k labels.
The work proposes FoMo: forward-noising a reference image to a sampled timestep and then denoising it independently, using that forking moment as a pointwise perceptual-distance label, so training data can be generated with no human annotation, and trains reference-based IQA metrics with a RankNet-style global ranking objective; across four benchmarks and seven backbones it achieves the best average performance, with LPIPS-Alex reaching 0.733 SROCC on PIPAL versus 0.577 for the same architecture trained on human-annotated KADID-10k labels.
The work proposes FoMo: forward-noising a reference image to a sampled timestep and then denoising it independently, using that forking moment as a pointwise perceptual-distance label, so training data can be generated with no human annotation, and trains reference-based IQA metrics with a RankNet-style global ranking objective; across four benchmarks and seven backbones it achieves the best average performance, with LPIPS-Alex reaching 0.733 SROCC on PIPAL versus 0.577 for the same architecture trained on human-annotated KADID-10k labels.
arXiv The work introduces Tactile-JEPA, a self-supervised pretraining method for distributed tactile sensors (e-skins) that samples masks at local and global scales over the taxel connectivity graph and predicts the embeddings of masked taxels; across magnetic and piezoresistive sensors and three datasets (Sparsh-skin, Tactile socks, DECO-50) it reduces force-estimation RMSE by 6.3% and in-hand pose RMSE by 20.8% over the strongest prior baseline, with consistent gains in action classification, object classification, and tactile-conditioned policy learning.
The work introduces Tactile-JEPA, a self-supervised pretraining method for distributed tactile sensors (e-skins) that samples masks at local and global scales over the taxel connectivity graph and predicts the embeddings of masked taxels; across magnetic and piezoresistive sensors and three datasets (Sparsh-skin, Tactile socks, DECO-50) it reduces force-estimation RMSE by 6.3% and in-hand pose RMSE by 20.8% over the strongest prior baseline, with consistent gains in action classification, object classification, and tactile-conditioned policy learning.
The work introduces Tactile-JEPA, a self-supervised pretraining method for distributed tactile sensors (e-skins) that samples masks at local and global scales over the taxel connectivity graph and predicts the embeddings of masked taxels; across magnetic and piezoresistive sensors and three datasets (Sparsh-skin, Tactile socks, DECO-50) it reduces force-estimation RMSE by 6.3% and in-hand pose RMSE by 20.8% over the strongest prior baseline, with consistent gains in action classification, object classification, and tactile-conditioned policy learning.
The work introduces Tactile-JEPA, a self-supervised pretraining method for distributed tactile sensors (e-skins) that samples masks at local and global scales over the taxel connectivity graph and predicts the embeddings of masked taxels; across magnetic and piezoresistive sensors and three datasets (Sparsh-skin, Tactile socks, DECO-50) it reduces force-estimation RMSE by 6.3% and in-hand pose RMSE by 20.8% over the strongest prior baseline, with consistent gains in action classification, object classification, and tactile-conditioned policy learning.
IEEE Spectrum This short piece presents 'Poetry for Engineers: The UI Designer's Dream,' a poem by Ralph Earle, a former technical editor and later senior software engineer at the IBM Software Lab, in which writing code and building the internet pixel by pixel is framed as translating dry code into surpassing beauty, and the speaker dreams of translating holy visions into a communication protocol of universal wonderment and letting shreds of light fall like diamonds on the endless mountains of the Web; an accompanying biography notes that Earle joined the IBM Software Lab in Raleigh, N.C., in 1990, retired as a senior software engineer in 2017, coauthored Enterprise Computing with Objects, and after retirement published two poetry collections and was nominated for the Pushcart Prize three times.
This short piece presents 'Poetry for Engineers: The UI Designer's Dream,' a poem by Ralph Earle, a former technical editor and later senior software engineer at the IBM Software Lab, in which writing code and building the internet pixel by pixel is framed as translating dry code into surpassing beauty, and the speaker dreams of translating holy visions into a communication protocol of universal wonderment and letting shreds of light fall like diamonds on the endless mountains of the Web; an accompanying biography notes that Earle joined the IBM Software Lab in Raleigh, N.C., in 1990, retired as a senior software engineer in 2017, coauthored Enterprise Computing with Objects, and after retirement published two poetry collections and was nominated for the Pushcart Prize three times.
This short piece presents 'Poetry for Engineers: The UI Designer's Dream,' a poem by Ralph Earle, a former technical editor and later senior software engineer at the IBM Software Lab, in which writing code and building the internet pixel by pixel is framed as translating dry code into surpassing beauty, and the speaker dreams of translating holy visions into a communication protocol of universal wonderment and letting shreds of light fall like diamonds on the endless mountains of the Web; an accompanying biography notes that Earle joined the IBM Software Lab in Raleigh, N.C., in 1990, retired as a senior software engineer in 2017, coauthored Enterprise Computing with Objects, and after retirement published two poetry collections and was nominated for the Pushcart Prize three times.
This short piece presents 'Poetry for Engineers: The UI Designer's Dream,' a poem by Ralph Earle, a former technical editor and later senior software engineer at the IBM Software Lab, in which writing code and building the internet pixel by pixel is framed as translating dry code into surpassing beauty, and the speaker dreams of translating holy visions into a communication protocol of universal wonderment and letting shreds of light fall like diamonds on the endless mountains of the Web; an accompanying biography notes that Earle joined the IBM Software Lab in Raleigh, N.C., in 1990, retired as a senior software engineer in 2017, coauthored Enterprise Computing with Objects, and after retirement published two poetry collections and was nominated for the Pushcart Prize three times.
bioRxiv The study presents a single-nucleus multi-omic atlas comprising 459,856 transcriptomic and chromatin accessibility profiles from 21 adult human tissues and four donors, including paired measurements from 160,688 nuclei, resolving nine cell lineages, 61 broad cell types and 313 subclusters, identifying 1,085,062 candidate cis-regulatory elements (including 161,270 novel elements absent from ENCODE), and using the dataset to train sequence-to-function models that predict chromatin-accessibility effects for 548,656 fine-mapped variants, identifying 18,133 high-effect variants including 1,120 broadly active variants.
The study presents a single-nucleus multi-omic atlas comprising 459,856 transcriptomic and chromatin accessibility profiles from 21 adult human tissues and four donors, including paired measurements from 160,688 nuclei, resolving nine cell lineages, 61 broad cell types and 313 subclusters, identifying 1,085,062 candidate cis-regulatory elements (including 161,270 novel elements absent from ENCODE), and using the dataset to train sequence-to-function models that predict chromatin-accessibility effects for 548,656 fine-mapped variants, identifying 18,133 high-effect variants including 1,120 broadly active variants.
The study presents a single-nucleus multi-omic atlas comprising 459,856 transcriptomic and chromatin accessibility profiles from 21 adult human tissues and four donors, including paired measurements from 160,688 nuclei, resolving nine cell lineages, 61 broad cell types and 313 subclusters, identifying 1,085,062 candidate cis-regulatory elements (including 161,270 novel elements absent from ENCODE), and using the dataset to train sequence-to-function models that predict chromatin-accessibility effects for 548,656 fine-mapped variants, identifying 18,133 high-effect variants including 1,120 broadly active variants.
The study presents a single-nucleus multi-omic atlas comprising 459,856 transcriptomic and chromatin accessibility profiles from 21 adult human tissues and four donors, including paired measurements from 160,688 nuclei, resolving nine cell lineages, 61 broad cell types and 313 subclusters, identifying 1,085,062 candidate cis-regulatory elements (including 161,270 novel elements absent from ENCODE), and using the dataset to train sequence-to-function models that predict chromatin-accessibility effects for 548,656 fine-mapped variants, identifying 18,133 high-effect variants including 1,120 broadly active variants.
发表出处待核验 This PhD thesis addresses how causal machine learning can be translated into real-world clinical practice along three lines: for binary treatment effect estimation in clinical contexts it identifies eight key decision points that directly influence treatment effect estimates and proposes a structured framework for designing causal models in clinical settings; for the difficulty of verifying counterfactual predictions it proposes a novel method that grounds average model predictions in a group-level quantity that can be empirically verified using randomized or real-world clinical trial data, and shows that a model with a lower counterfactual error bound can be constructed and that information bottleneck regularization improves treatment effect estimation; for continuous causal machine learn
This PhD thesis addresses how causal machine learning can be translated into real-world clinical practice along three lines: for binary treatment effect estimation in clinical contexts it identifies eight key decision points that directly influence treatment effect estimates and proposes a structured framework for designing causal models in clinical settings; for the difficulty of verifying counterfactual predictions it proposes a novel method that grounds average model predictions in a group-level quantity that can be empirically verified using randomized or real-world clinical trial data, and shows that a model with a lower counterfactual error bound can be constructed and that information bottleneck regularization improves treatment effect estimation; for continuous causal machine learn
This PhD thesis addresses how causal machine learning can be translated into real-world clinical practice along three lines: for binary treatment effect estimation in clinical contexts it identifies eight key decision points that directly influence treatment effect estimates and proposes a structured framework for designing causal models in clinical settings; for the difficulty of verifying counterfactual predictions it proposes a novel method that grounds average model predictions in a group-level quantity that can be empirically verified using randomized or real-world clinical trial data, and shows that a model with a lower counterfactual error bound can be constructed and that information bottleneck regularization improves treatment effect estimation; for continuous causal machine learn
This PhD thesis addresses how causal machine learning can be translated into real-world clinical practice along three lines: for binary treatment effect estimation in clinical contexts it identifies eight key decision points that directly influence treatment effect estimates and proposes a structured framework for designing causal models in clinical settings; for the difficulty of verifying counterfactual predictions it proposes a novel method that grounds average model predictions in a group-level quantity that can be empirically verified using randomized or real-world clinical trial data, and shows that a model with a lower counterfactual error bound can be constructed and that information bottleneck regularization improves treatment effect estimation; for continuous causal machine learn
International Journal of Modern Science and Research Technology This narrative review searched peer-reviewed GenAI studies published through September 05, 2026 alongside foundational cognitive-offloading research and found that evidence does not support a simple positive or negative effect of GenAI on human cognition: GenAI-supported offloading of functions such as memory retrieval, processing, and learning can reduce required cognitive capacity and support high performance across a variety of tasks, yet the same offloading can in some cases reduce the quality of subsequent encoding, and recent GenAI studies show a consistent performance-learning dissociation in which gains from support available at the time typically carry over to later unsupported performance only slightly, not at all, or even negatively.
This narrative review searched peer-reviewed GenAI studies published through September 05, 2026 alongside foundational cognitive-offloading research and found that evidence does not support a simple positive or negative effect of GenAI on human cognition: GenAI-supported offloading of functions such as memory retrieval, processing, and learning can reduce required cognitive capacity and support high performance across a variety of tasks, yet the same offloading can in some cases reduce the quality of subsequent encoding, and recent GenAI studies show a consistent performance-learning dissociation in which gains from support available at the time typically carry over to later unsupported performance only slightly, not at all, or even negatively.
This narrative review searched peer-reviewed GenAI studies published through September 05, 2026 alongside foundational cognitive-offloading research and found that evidence does not support a simple positive or negative effect of GenAI on human cognition: GenAI-supported offloading of functions such as memory retrieval, processing, and learning can reduce required cognitive capacity and support high performance across a variety of tasks, yet the same offloading can in some cases reduce the quality of subsequent encoding, and recent GenAI studies show a consistent performance-learning dissociation in which gains from support available at the time typically carry over to later unsupported performance only slightly, not at all, or even negatively.
This narrative review searched peer-reviewed GenAI studies published through September 05, 2026 alongside foundational cognitive-offloading research and found that evidence does not support a simple positive or negative effect of GenAI on human cognition: GenAI-supported offloading of functions such as memory retrieval, processing, and learning can reduce required cognitive capacity and support high performance across a variety of tasks, yet the same offloading can in some cases reduce the quality of subsequent encoding, and recent GenAI studies show a consistent performance-learning dissociation in which gains from support available at the time typically carry over to later unsupported performance only slightly, not at all, or even negatively.
发表出处待核验 This doctoral thesis, conducted within the EU-funded I-CARE4OLD project, first maps the definition and use of non-pharmacological interventions (NPIs)—a systematic review shows highly heterogeneous and predominantly negatively framed definitions, and data from six countries show underuse of interventions with proven benefits (physical therapy, occupational therapy, psychosocial interventions, social participation) alongside relatively frequent use of physical restraints in some countries—and second, using longitudinal real-world data from the Netherlands and Ontario, Canada, with propensity-based methods, generalized estimating equations, and target trial emulation with causal machine learning, finds that moderate and higher levels of physical activity are associated with beneficial effect
This doctoral thesis, conducted within the EU-funded I-CARE4OLD project, first maps the definition and use of non-pharmacological interventions (NPIs)—a systematic review shows highly heterogeneous and predominantly negatively framed definitions, and data from six countries show underuse of interventions with proven benefits (physical therapy, occupational therapy, psychosocial interventions, social participation) alongside relatively frequent use of physical restraints in some countries—and second, using longitudinal real-world data from the Netherlands and Ontario, Canada, with propensity-based methods, generalized estimating equations, and target trial emulation with causal machine learning, finds that moderate and higher levels of physical activity are associated with beneficial effect
This doctoral thesis, conducted within the EU-funded I-CARE4OLD project, first maps the definition and use of non-pharmacological interventions (NPIs)—a systematic review shows highly heterogeneous and predominantly negatively framed definitions, and data from six countries show underuse of interventions with proven benefits (physical therapy, occupational therapy, psychosocial interventions, social participation) alongside relatively frequent use of physical restraints in some countries—and second, using longitudinal real-world data from the Netherlands and Ontario, Canada, with propensity-based methods, generalized estimating equations, and target trial emulation with causal machine learning, finds that moderate and higher levels of physical activity are associated with beneficial effect
This doctoral thesis, conducted within the EU-funded I-CARE4OLD project, first maps the definition and use of non-pharmacological interventions (NPIs)—a systematic review shows highly heterogeneous and predominantly negatively framed definitions, and data from six countries show underuse of interventions with proven benefits (physical therapy, occupational therapy, psychosocial interventions, social participation) alongside relatively frequent use of physical restraints in some countries—and second, using longitudinal real-world data from the Netherlands and Ontario, Canada, with propensity-based methods, generalized estimating equations, and target trial emulation with causal machine learning, finds that moderate and higher levels of physical activity are associated with beneficial effect
IJIS - Indonesian Journal On Information System The study collected user reviews of the EV Charging/SPKLU feature in PLN Mobile from Google Play Store, labeled sentiment using ratings as weak labels, applied algorithm-specific class weighting on the training data to address class imbalance, and compared seven classifiers—Logistic Regression, SVM, Naive Bayes, MLP, RNN, BiLSTM, and DistilBERT—finding that MLP achieved the highest validation macro F1 (67.80%) and was selected for final testing, where it reached 89.16% accuracy and 59.42% macro F1 on 286 test instances, indicating high overall classification performance with uneven performance across sentiment classes.
The study collected user reviews of the EV Charging/SPKLU feature in PLN Mobile from Google Play Store, labeled sentiment using ratings as weak labels, applied algorithm-specific class weighting on the training data to address class imbalance, and compared seven classifiers—Logistic Regression, SVM, Naive Bayes, MLP, RNN, BiLSTM, and DistilBERT—finding that MLP achieved the highest validation macro F1 (67.80%) and was selected for final testing, where it reached 89.16% accuracy and 59.42% macro F1 on 286 test instances, indicating high overall classification performance with uneven performance across sentiment classes.
The study collected user reviews of the EV Charging/SPKLU feature in PLN Mobile from Google Play Store, labeled sentiment using ratings as weak labels, applied algorithm-specific class weighting on the training data to address class imbalance, and compared seven classifiers—Logistic Regression, SVM, Naive Bayes, MLP, RNN, BiLSTM, and DistilBERT—finding that MLP achieved the highest validation macro F1 (67.80%) and was selected for final testing, where it reached 89.16% accuracy and 59.42% macro F1 on 286 test instances, indicating high overall classification performance with uneven performance across sentiment classes.
The study collected user reviews of the EV Charging/SPKLU feature in PLN Mobile from Google Play Store, labeled sentiment using ratings as weak labels, applied algorithm-specific class weighting on the training data to address class imbalance, and compared seven classifiers—Logistic Regression, SVM, Naive Bayes, MLP, RNN, BiLSTM, and DistilBERT—finding that MLP achieved the highest validation macro F1 (67.80%) and was selected for final testing, where it reached 89.16% accuracy and 59.42% macro F1 on 286 test instances, indicating high overall classification performance with uneven performance across sentiment classes.
JURNAL PESISIR DAN LAUT TROPIS Using MAWS observations, this study evaluated the bias in BMKG Ina-Flows sea surface temperature forecasts and compared three machine-learning bias-correction algorithms (SVR, LSTM, Bi-LSTM) at Karimun Jawa and Bira, finding a warm bias at both sites, with LSTM best at Karimun Jawa (RMSE 0.207, MAE 0.182, MBE -0.171, a 29.83% RMSE reduction) and Bi-LSTM best at Bira (RMSE 0.218, MAE 0.173, MBE 0.011, a 26.10% RMSE reduction), indicating that correction effectiveness is location-dependent and algorithm choice should rest on local validation.
Using MAWS observations, this study evaluated the bias in BMKG Ina-Flows sea surface temperature forecasts and compared three machine-learning bias-correction algorithms (SVR, LSTM, Bi-LSTM) at Karimun Jawa and Bira, finding a warm bias at both sites, with LSTM best at Karimun Jawa (RMSE 0.207, MAE 0.182, MBE -0.171, a 29.83% RMSE reduction) and Bi-LSTM best at Bira (RMSE 0.218, MAE 0.173, MBE 0.011, a 26.10% RMSE reduction), indicating that correction effectiveness is location-dependent and algorithm choice should rest on local validation.
Using MAWS observations, this study evaluated the bias in BMKG Ina-Flows sea surface temperature forecasts and compared three machine-learning bias-correction algorithms (SVR, LSTM, Bi-LSTM) at Karimun Jawa and Bira, finding a warm bias at both sites, with LSTM best at Karimun Jawa (RMSE 0.207, MAE 0.182, MBE -0.171, a 29.83% RMSE reduction) and Bi-LSTM best at Bira (RMSE 0.218, MAE 0.173, MBE 0.011, a 26.10% RMSE reduction), indicating that correction effectiveness is location-dependent and algorithm choice should rest on local validation.
Using MAWS observations, this study evaluated the bias in BMKG Ina-Flows sea surface temperature forecasts and compared three machine-learning bias-correction algorithms (SVR, LSTM, Bi-LSTM) at Karimun Jawa and Bira, finding a warm bias at both sites, with LSTM best at Karimun Jawa (RMSE 0.207, MAE 0.182, MBE -0.171, a 29.83% RMSE reduction) and Bi-LSTM best at Bira (RMSE 0.218, MAE 0.173, MBE 0.011, a 26.10% RMSE reduction), indicating that correction effectiveness is location-dependent and algorithm choice should rest on local validation.
arXiv The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
arXiv The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a setting where the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under fresh, previously undisclosed requirements and network isolation it obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models showing clear same/cross-model separation and finite-chal
The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a setting where the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under fresh, previously undisclosed requirements and network isolation it obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models showing clear same/cross-model separation and finite-chal
The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a setting where the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under fresh, previously undisclosed requirements and network isolation it obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models showing clear same/cross-model separation and finite-chal
The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a setting where the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under fresh, previously undisclosed requirements and network isolation it obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models showing clear same/cross-model separation and finite-chal
arXiv The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs, introducing a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism; evaluation shows the benchmark is far from saturated, with the best model Gemini-3.1-pro-preview reaching only 43.5%, substitution-resistant evaluation yielding a top score of 30.6% for GPT-5.4 while Gemini-3.1-pro-preview drops to 17.6% below the 20% random baseline, and a dual-mode protocol showing consistent accuracy gains from proof-sketch access.
The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs, introducing a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism; evaluation shows the benchmark is far from saturated, with the best model Gemini-3.1-pro-preview reaching only 43.5%, substitution-resistant evaluation yielding a top score of 30.6% for GPT-5.4 while Gemini-3.1-pro-preview drops to 17.6% below the 20% random baseline, and a dual-mode protocol showing consistent accuracy gains from proof-sketch access.
The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs, introducing a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism; evaluation shows the benchmark is far from saturated, with the best model Gemini-3.1-pro-preview reaching only 43.5%, substitution-resistant evaluation yielding a top score of 30.6% for GPT-5.4 while Gemini-3.1-pro-preview drops to 17.6% below the 20% random baseline, and a dual-mode protocol showing consistent accuracy gains from proof-sketch access.
The work presents LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs, introducing a thirteen-category logical taxonomy of theorem types, a proof-sketch-guided distractor pipeline, and a substitution-resistant mechanism; evaluation shows the benchmark is far from saturated, with the best model Gemini-3.1-pro-preview reaching only 43.5%, substitution-resistant evaluation yielding a top score of 30.6% for GPT-5.4 while Gemini-3.1-pro-preview drops to 17.6% below the 20% random baseline, and a dual-mode protocol showing consistent accuracy gains from proof-sketch access.
arXiv The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod