Only content delivered through the publication boundary on this date is included.
arXiv This work demonstrates on a live O-RAN testbed that two autonomous agents with individually correct objectives—one protecting a latency SLA and one maximizing utilization for energy efficiency—jointly drive recurring opposing excursions of the shared resource partition, and presents and proves a lightweight arbitration layer, AURA, whose three admission checks (feasibility invariants, per-variable dwell time, deadband) reduce shared-state excursion amplitude from 8.4 to 0.4 PRB and cross-slice throughput starvation from 40–55% to 0.3%, while leaving the protected slice's own latency compliance unchanged.
This work demonstrates on a live O-RAN testbed that two autonomous agents with individually correct objectives—one protecting a latency SLA and one maximizing utilization for energy efficiency—jointly drive recurring opposing excursions of the shared resource partition, and presents and proves a lightweight arbitration layer, AURA, whose three admission checks (feasibility invariants, per-variable dwell time, deadband) reduce shared-state excursion amplitude from 8.4 to 0.4 PRB and cross-slice throughput starvation from 40–55% to 0.3%, while leaving the protected slice's own latency compliance unchanged.
This work demonstrates on a live O-RAN testbed that two autonomous agents with individually correct objectives—one protecting a latency SLA and one maximizing utilization for energy efficiency—jointly drive recurring opposing excursions of the shared resource partition, and presents and proves a lightweight arbitration layer, AURA, whose three admission checks (feasibility invariants, per-variable dwell time, deadband) reduce shared-state excursion amplitude from 8.4 to 0.4 PRB and cross-slice throughput starvation from 40–55% to 0.3%, while leaving the protected slice's own latency compliance unchanged.
This work demonstrates on a live O-RAN testbed that two autonomous agents with individually correct objectives—one protecting a latency SLA and one maximizing utilization for energy efficiency—jointly drive recurring opposing excursions of the shared resource partition, and presents and proves a lightweight arbitration layer, AURA, whose three admission checks (feasibility invariants, per-variable dwell time, deadband) reduce shared-state excursion amplitude from 8.4 to 0.4 PRB and cross-slice throughput starvation from 40–55% to 0.3%, while leaving the protected slice's own latency compliance unchanged.
arXiv EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.
EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.
EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.
EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.
arXiv The work proposes that running tool calls explicitly report their own progress (a fraction of work remaining or an accurate signal that the end is near), and through a census of four public agent corpora, a harness that recovers the signal without changing what the agent sees, a comparison against four published predictors, and an end-to-end evaluation plugged into vLLM through a few small hints, finds that the reported progress is between several times and an order of magnitude more accurate than the best published predictors at KV cache decision points and stays accurate when the environment changes, cutting p90 time to first token after a tool call by 20.7% (HBM only) and 20.8% (HBM + DRAM) against LRU, close to an oracle.
The work proposes that running tool calls explicitly report their own progress (a fraction of work remaining or an accurate signal that the end is near), and through a census of four public agent corpora, a harness that recovers the signal without changing what the agent sees, a comparison against four published predictors, and an end-to-end evaluation plugged into vLLM through a few small hints, finds that the reported progress is between several times and an order of magnitude more accurate than the best published predictors at KV cache decision points and stays accurate when the environment changes, cutting p90 time to first token after a tool call by 20.7% (HBM only) and 20.8% (HBM + DRAM) against LRU, close to an oracle.
The work proposes that running tool calls explicitly report their own progress (a fraction of work remaining or an accurate signal that the end is near), and through a census of four public agent corpora, a harness that recovers the signal without changing what the agent sees, a comparison against four published predictors, and an end-to-end evaluation plugged into vLLM through a few small hints, finds that the reported progress is between several times and an order of magnitude more accurate than the best published predictors at KV cache decision points and stays accurate when the environment changes, cutting p90 time to first token after a tool call by 20.7% (HBM only) and 20.8% (HBM + DRAM) against LRU, close to an oracle.
The work proposes that running tool calls explicitly report their own progress (a fraction of work remaining or an accurate signal that the end is near), and through a census of four public agent corpora, a harness that recovers the signal without changing what the agent sees, a comparison against four published predictors, and an end-to-end evaluation plugged into vLLM through a few small hints, finds that the reported progress is between several times and an order of magnitude more accurate than the best published predictors at KV cache decision points and stays accurate when the environment changes, cutting p90 time to first token after a tool call by 20.7% (HBM only) and 20.8% (HBM + DRAM) against LRU, close to an oracle.
arXiv The work proposes an "infinite-parameter LLM" architecture in which a compact hypernetwork encodes run-time data (facts, instructions, demonstrations) into a low-dimensional latent code, that code generates a low-rank additive modulation of a shared base feed-forward network so each token's "expert" is generated rather than drawn from a stored bank, and a Bayesian belief over the latent code is carried and updated online so the effective weight keeps evolving within a session; the authors also specify an evaluation protocol that pits carrying knowledge in the weights against carrying it in the prompt at matched budget.
The work proposes an "infinite-parameter LLM" architecture in which a compact hypernetwork encodes run-time data (facts, instructions, demonstrations) into a low-dimensional latent code, that code generates a low-rank additive modulation of a shared base feed-forward network so each token's "expert" is generated rather than drawn from a stored bank, and a Bayesian belief over the latent code is carried and updated online so the effective weight keeps evolving within a session; the authors also specify an evaluation protocol that pits carrying knowledge in the weights against carrying it in the prompt at matched budget.
The work proposes an "infinite-parameter LLM" architecture in which a compact hypernetwork encodes run-time data (facts, instructions, demonstrations) into a low-dimensional latent code, that code generates a low-rank additive modulation of a shared base feed-forward network so each token's "expert" is generated rather than drawn from a stored bank, and a Bayesian belief over the latent code is carried and updated online so the effective weight keeps evolving within a session; the authors also specify an evaluation protocol that pits carrying knowledge in the weights against carrying it in the prompt at matched budget.
The work proposes an "infinite-parameter LLM" architecture in which a compact hypernetwork encodes run-time data (facts, instructions, demonstrations) into a low-dimensional latent code, that code generates a low-rank additive modulation of a shared base feed-forward network so each token's "expert" is generated rather than drawn from a stored bank, and a Bayesian belief over the latent code is carried and updated online so the effective weight keeps evolving within a session; the authors also specify an evaluation protocol that pits carrying knowledge in the weights against carrying it in the prompt at matched budget.
arXiv The paper establishes a degree-independent spectral independence bound for log-concave Holant problems on simple graphs, yielding relaxation-time bounds for Glauber dynamics of O_λ(m) for the monomer–dimer model at activity λ, O_{b,λ}(m) for b-matchings at fugacity λ>0, and O(bm) for uniform b-matchings, where m is the number of edges.
The paper establishes a degree-independent spectral independence bound for log-concave Holant problems on simple graphs, yielding relaxation-time bounds for Glauber dynamics of O_λ(m) for the monomer–dimer model at activity λ, O_{b,λ}(m) for b-matchings at fugacity λ>0, and O(bm) for uniform b-matchings, where m is the number of edges.
The paper establishes a degree-independent spectral independence bound for log-concave Holant problems on simple graphs, yielding relaxation-time bounds for Glauber dynamics of O_λ(m) for the monomer–dimer model at activity λ, O_{b,λ}(m) for b-matchings at fugacity λ>0, and O(bm) for uniform b-matchings, where m is the number of edges.
The paper establishes a degree-independent spectral independence bound for log-concave Holant problems on simple graphs, yielding relaxation-time bounds for Glauber dynamics of O_λ(m) for the monomer–dimer model at activity λ, O_{b,λ}(m) for b-matchings at fugacity λ>0, and O(bm) for uniform b-matchings, where m is the number of edges.
arXiv Across four vision-language models (Qwen3-VL-2B, Qwen3-VL-8B, Molmo2-7B, Llava-Next-34B), the work identifies attention heads causally necessary for OCR and shows they are general-purpose verbalization heads; summing their output-value matrices yields a verbalization lens that decodes image-token hidden states into interpretable semantic words at any layer, including layer 0, and whose pseudo-inverse provides concept vectors for editing image representations (e.g., replacing a tractor with a revolver), supporting the view that image representations are aligned with language space from early layers.
Across four vision-language models (Qwen3-VL-2B, Qwen3-VL-8B, Molmo2-7B, Llava-Next-34B), the work identifies attention heads causally necessary for OCR and shows they are general-purpose verbalization heads; summing their output-value matrices yields a verbalization lens that decodes image-token hidden states into interpretable semantic words at any layer, including layer 0, and whose pseudo-inverse provides concept vectors for editing image representations (e.g., replacing a tractor with a revolver), supporting the view that image representations are aligned with language space from early layers.
Across four vision-language models (Qwen3-VL-2B, Qwen3-VL-8B, Molmo2-7B, Llava-Next-34B), the work identifies attention heads causally necessary for OCR and shows they are general-purpose verbalization heads; summing their output-value matrices yields a verbalization lens that decodes image-token hidden states into interpretable semantic words at any layer, including layer 0, and whose pseudo-inverse provides concept vectors for editing image representations (e.g., replacing a tractor with a revolver), supporting the view that image representations are aligned with language space from early layers.
Across four vision-language models (Qwen3-VL-2B, Qwen3-VL-8B, Molmo2-7B, Llava-Next-34B), the work identifies attention heads causally necessary for OCR and shows they are general-purpose verbalization heads; summing their output-value matrices yields a verbalization lens that decodes image-token hidden states into interpretable semantic words at any layer, including layer 0, and whose pseudo-inverse provides concept vectors for editing image representations (e.g., replacing a tractor with a revolver), supporting the view that image representations are aligned with language space from early layers.
arXiv The paper defines and formalizes Compositional Policy Violations (CPVs), a governance failure mode in which every step of an agentic workflow passes its own local check while the composed execution violates the governing policy, and it offers a four-part taxonomy (Authority Creep, Threshold Laundering, Cumulative Sum Violation, Context Collapse), argues that the correct repair topology is dictated by where the guarded quantity mutates, and proposes a provenance-aware runtime architecture that evaluates policies over complete execution traces and recomputes guarded quantities from raw provenance rather than the pipeline's derived representation.
The paper defines and formalizes Compositional Policy Violations (CPVs), a governance failure mode in which every step of an agentic workflow passes its own local check while the composed execution violates the governing policy, and it offers a four-part taxonomy (Authority Creep, Threshold Laundering, Cumulative Sum Violation, Context Collapse), argues that the correct repair topology is dictated by where the guarded quantity mutates, and proposes a provenance-aware runtime architecture that evaluates policies over complete execution traces and recomputes guarded quantities from raw provenance rather than the pipeline's derived representation.
The paper defines and formalizes Compositional Policy Violations (CPVs), a governance failure mode in which every step of an agentic workflow passes its own local check while the composed execution violates the governing policy, and it offers a four-part taxonomy (Authority Creep, Threshold Laundering, Cumulative Sum Violation, Context Collapse), argues that the correct repair topology is dictated by where the guarded quantity mutates, and proposes a provenance-aware runtime architecture that evaluates policies over complete execution traces and recomputes guarded quantities from raw provenance rather than the pipeline's derived representation.
The paper defines and formalizes Compositional Policy Violations (CPVs), a governance failure mode in which every step of an agentic workflow passes its own local check while the composed execution violates the governing policy, and it offers a four-part taxonomy (Authority Creep, Threshold Laundering, Cumulative Sum Violation, Context Collapse), argues that the correct repair topology is dictated by where the guarded quantity mutates, and proposes a provenance-aware runtime architecture that evaluates policies over complete execution traces and recomputes guarded quantities from raw provenance rather than the pipeline's derived representation.
Terence Tao blog RSS This guest post by the editors of the communal blog Proofs and Prompts describes why the blog was started and what it aims to do: as AI became a central topic of conversation in the mathematical community, the editors created a collective space for mathematicians who lack a natural platform, especially PhD students, and reports that about a month after launch contributors ranged from PhD students and undergrads to Fields medallists and hobbyists, with posts ranging from a call for a general moratorium on AI to the view that AI plays too small a role in mathematics, while noting that contributors remain concentrated in Western Europe and North America, few have experience developing or evaluating LLMs, and only a handful of the forty-or-so published posts were written by women, and inviting
This guest post by the editors of the communal blog Proofs and Prompts describes why the blog was started and what it aims to do: as AI became a central topic of conversation in the mathematical community, the editors created a collective space for mathematicians who lack a natural platform, especially PhD students, and reports that about a month after launch contributors ranged from PhD students and undergrads to Fields medallists and hobbyists, with posts ranging from a call for a general moratorium on AI to the view that AI plays too small a role in mathematics, while noting that contributors remain concentrated in Western Europe and North America, few have experience developing or evaluating LLMs, and only a handful of the forty-or-so published posts were written by women, and inviting
This guest post by the editors of the communal blog Proofs and Prompts describes why the blog was started and what it aims to do: as AI became a central topic of conversation in the mathematical community, the editors created a collective space for mathematicians who lack a natural platform, especially PhD students, and reports that about a month after launch contributors ranged from PhD students and undergrads to Fields medallists and hobbyists, with posts ranging from a call for a general moratorium on AI to the view that AI plays too small a role in mathematics, while noting that contributors remain concentrated in Western Europe and North America, few have experience developing or evaluating LLMs, and only a handful of the forty-or-so published posts were written by women, and inviting
This guest post by the editors of the communal blog Proofs and Prompts describes why the blog was started and what it aims to do: as AI became a central topic of conversation in the mathematical community, the editors created a collective space for mathematicians who lack a natural platform, especially PhD students, and reports that about a month after launch contributors ranged from PhD students and undergrads to Fields medallists and hobbyists, with posts ranging from a call for a general moratorium on AI to the view that AI plays too small a role in mathematics, while noting that contributors remain concentrated in Western Europe and North America, few have experience developing or evaluating LLMs, and only a handful of the forty-or-so published posts were written by women, and inviting
arXiv This work introduces the ProgramDistill benchmark and a fully automated mine–craft–patch pipeline that factorizes 26 interactive web applications into 1,975 replay-verified behaviors and constructs 4,063 tasks, letting coding agents infer and restore missing functionality by interacting with a reference application whose source is hidden; across nine frontier agents, the best full-application reconstruction success is 49.2%, and partial-application reconstruction success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.
This work introduces the ProgramDistill benchmark and a fully automated mine–craft–patch pipeline that factorizes 26 interactive web applications into 1,975 replay-verified behaviors and constructs 4,063 tasks, letting coding agents infer and restore missing functionality by interacting with a reference application whose source is hidden; across nine frontier agents, the best full-application reconstruction success is 49.2%, and partial-application reconstruction success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.
This work introduces the ProgramDistill benchmark and a fully automated mine–craft–patch pipeline that factorizes 26 interactive web applications into 1,975 replay-verified behaviors and constructs 4,063 tasks, letting coding agents infer and restore missing functionality by interacting with a reference application whose source is hidden; across nine frontier agents, the best full-application reconstruction success is 49.2%, and partial-application reconstruction success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.
This work introduces the ProgramDistill benchmark and a fully automated mine–craft–patch pipeline that factorizes 26 interactive web applications into 1,975 replay-verified behaviors and constructs 4,063 tasks, letting coding agents infer and restore missing functionality by interacting with a reference application whose source is hidden; across nine frontier agents, the best full-application reconstruction success is 49.2%, and partial-application reconstruction success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.
arXiv This work studies Transient Multiagent Pathfinding, in which agents must be routed without collisions and disappear upon reaching their destinations, and shows that the problem is fixed-parameter tractable in the combined parameter k+ζ, where k is the number of agents and ζ=L−λ is the gap between the sequential-routing upper bound L=1+Σdist(si,ti) and the target makespan λ, running in 2^{O(k²ζ)}·n^{O(1)} time, complemented by matching lower bounds (W[1]-hardness for k alone, W[1]-hardness for ζ alone when terminals need not be distinct, and no polynomial kernel for k+ζ) and by fixed-parameter tractability in ζ alone when all terminals are pairwise distinct, running in 2^{O(ζ³)}·n^{O(1)} time.
This work studies Transient Multiagent Pathfinding, in which agents must be routed without collisions and disappear upon reaching their destinations, and shows that the problem is fixed-parameter tractable in the combined parameter k+ζ, where k is the number of agents and ζ=L−λ is the gap between the sequential-routing upper bound L=1+Σdist(si,ti) and the target makespan λ, running in 2^{O(k²ζ)}·n^{O(1)} time, complemented by matching lower bounds (W[1]-hardness for k alone, W[1]-hardness for ζ alone when terminals need not be distinct, and no polynomial kernel for k+ζ) and by fixed-parameter tractability in ζ alone when all terminals are pairwise distinct, running in 2^{O(ζ³)}·n^{O(1)} time.
This work studies Transient Multiagent Pathfinding, in which agents must be routed without collisions and disappear upon reaching their destinations, and shows that the problem is fixed-parameter tractable in the combined parameter k+ζ, where k is the number of agents and ζ=L−λ is the gap between the sequential-routing upper bound L=1+Σdist(si,ti) and the target makespan λ, running in 2^{O(k²ζ)}·n^{O(1)} time, complemented by matching lower bounds (W[1]-hardness for k alone, W[1]-hardness for ζ alone when terminals need not be distinct, and no polynomial kernel for k+ζ) and by fixed-parameter tractability in ζ alone when all terminals are pairwise distinct, running in 2^{O(ζ³)}·n^{O(1)} time.
This work studies Transient Multiagent Pathfinding, in which agents must be routed without collisions and disappear upon reaching their destinations, and shows that the problem is fixed-parameter tractable in the combined parameter k+ζ, where k is the number of agents and ζ=L−λ is the gap between the sequential-routing upper bound L=1+Σdist(si,ti) and the target makespan λ, running in 2^{O(k²ζ)}·n^{O(1)} time, complemented by matching lower bounds (W[1]-hardness for k alone, W[1]-hardness for ζ alone when terminals need not be distinct, and no polynomial kernel for k+ζ) and by fixed-parameter tractability in ζ alone when all terminals are pairwise distinct, running in 2^{O(ζ³)}·n^{O(1)} time.
arXiv This work introduces AdaGeoVLN, a streaming vision-language navigation framework that couples VGGT geometry-foundation-model representations at depths 11, 17, and 23 to the first three Qwen3.5-4B decoder layers and retains historical VGGT global-attention KV states under a fixed budget according to instruction relevance, geometric confidence, and transition novelty, achieving 55.7%/51.4% and 54.1%/44.7% SR/SPL on R2R-CE and RxR-CE Val-Unseen with a single RGB stream and no additional navigation-specific external data, with ablations showing that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations and that bounded navigation-aware retention preserves navigation performance while reducing GFM-KV memory.
This work introduces AdaGeoVLN, a streaming vision-language navigation framework that couples VGGT geometry-foundation-model representations at depths 11, 17, and 23 to the first three Qwen3.5-4B decoder layers and retains historical VGGT global-attention KV states under a fixed budget according to instruction relevance, geometric confidence, and transition novelty, achieving 55.7%/51.4% and 54.1%/44.7% SR/SPL on R2R-CE and RxR-CE Val-Unseen with a single RGB stream and no additional navigation-specific external data, with ablations showing that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations and that bounded navigation-aware retention preserves navigation performance while reducing GFM-KV memory.
This work introduces AdaGeoVLN, a streaming vision-language navigation framework that couples VGGT geometry-foundation-model representations at depths 11, 17, and 23 to the first three Qwen3.5-4B decoder layers and retains historical VGGT global-attention KV states under a fixed budget according to instruction relevance, geometric confidence, and transition novelty, achieving 55.7%/51.4% and 54.1%/44.7% SR/SPL on R2R-CE and RxR-CE Val-Unseen with a single RGB stream and no additional navigation-specific external data, with ablations showing that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations and that bounded navigation-aware retention preserves navigation performance while reducing GFM-KV memory.
This work introduces AdaGeoVLN, a streaming vision-language navigation framework that couples VGGT geometry-foundation-model representations at depths 11, 17, and 23 to the first three Qwen3.5-4B decoder layers and retains historical VGGT global-attention KV states under a fixed budget according to instruction relevance, geometric confidence, and transition novelty, achieving 55.7%/51.4% and 54.1%/44.7% SR/SPL on R2R-CE and RxR-CE Val-Unseen with a single RGB stream and no additional navigation-specific external data, with ablations showing that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations and that bounded navigation-aware retention preserves navigation performance while reducing GFM-KV memory.
arXiv This work introduces CERA-MoA, an iterative reinforcement learning framework in which a dynamic router and multiple independent agent policies co-evolve: the router uses a predictive familiarity estimator built on mid-layer hidden states to assess each agent's semantic competence before generation, applies cumulative-threshold adaptive routing to activate a minimal agent subset, and proactively allocates targeted training samples based on agents' evolving competence, outperforming static-agent routing and fix-workflow fine-tuning baselines across diverse domains.
This work introduces CERA-MoA, an iterative reinforcement learning framework in which a dynamic router and multiple independent agent policies co-evolve: the router uses a predictive familiarity estimator built on mid-layer hidden states to assess each agent's semantic competence before generation, applies cumulative-threshold adaptive routing to activate a minimal agent subset, and proactively allocates targeted training samples based on agents' evolving competence, outperforming static-agent routing and fix-workflow fine-tuning baselines across diverse domains.
This work introduces CERA-MoA, an iterative reinforcement learning framework in which a dynamic router and multiple independent agent policies co-evolve: the router uses a predictive familiarity estimator built on mid-layer hidden states to assess each agent's semantic competence before generation, applies cumulative-threshold adaptive routing to activate a minimal agent subset, and proactively allocates targeted training samples based on agents' evolving competence, outperforming static-agent routing and fix-workflow fine-tuning baselines across diverse domains.
This work introduces CERA-MoA, an iterative reinforcement learning framework in which a dynamic router and multiple independent agent policies co-evolve: the router uses a predictive familiarity estimator built on mid-layer hidden states to assess each agent's semantic competence before generation, applies cumulative-threshold adaptive routing to activate a minimal agent subset, and proactively allocates targeted training samples based on agents' evolving competence, outperforming static-agent routing and fix-workflow fine-tuning baselines across diverse domains.
NVIDIA Research NVIDIA's first MLPerf Inference v6.1 preview submission of Vera Rubin NVL72 reports up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1, while GB300 NVL72 reached 99% scaling efficiency across 288 GPUs in four racks and software optimizations delivered up to 1.6x over v6.0.
NVIDIA's first MLPerf Inference v6.1 preview submission of Vera Rubin NVL72 reports up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1, while GB300 NVL72 reached 99% scaling efficiency across 288 GPUs in four racks and software optimizations delivered up to 1.6x over v6.0.
NVIDIA's first MLPerf Inference v6.1 preview submission of Vera Rubin NVL72 reports up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1, while GB300 NVL72 reached 99% scaling efficiency across 288 GPUs in four racks and software optimizations delivered up to 1.6x over v6.0.
NVIDIA's first MLPerf Inference v6.1 preview submission of Vera Rubin NVL72 reports up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1, while GB300 NVL72 reached 99% scaling efficiency across 288 GPUs in four racks and software optimizations delivered up to 1.6x over v6.0.
arXiv The work presents TeleAntiFraud 2.0, a refreshable Chinese call-audio benchmark organized as monthly frozen snapshots, built with a Mixed-Tree Anti-Fraud Generation Pipeline that turns online fraud-case abstracts into profile-grounded scenarios and expands them into fraud and near-domain lawful sibling dialogues sharing context and diverging only at label-bearing actions, rendered as role-matched speech; each frozen set contains 900 Chinese calls (600 fraud, 300 near-domain non-fraud), controlled text experiments show three classifiers reach perfect Macro-F1 against unrelated or ordinary negatives but drop to 0.65-0.68 with near-domain sibling negatives, and full-set audio and ASR+LLM evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity.
The work presents TeleAntiFraud 2.0, a refreshable Chinese call-audio benchmark organized as monthly frozen snapshots, built with a Mixed-Tree Anti-Fraud Generation Pipeline that turns online fraud-case abstracts into profile-grounded scenarios and expands them into fraud and near-domain lawful sibling dialogues sharing context and diverging only at label-bearing actions, rendered as role-matched speech; each frozen set contains 900 Chinese calls (600 fraud, 300 near-domain non-fraud), controlled text experiments show three classifiers reach perfect Macro-F1 against unrelated or ordinary negatives but drop to 0.65-0.68 with near-domain sibling negatives, and full-set audio and ASR+LLM evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity.
The work presents TeleAntiFraud 2.0, a refreshable Chinese call-audio benchmark organized as monthly frozen snapshots, built with a Mixed-Tree Anti-Fraud Generation Pipeline that turns online fraud-case abstracts into profile-grounded scenarios and expands them into fraud and near-domain lawful sibling dialogues sharing context and diverging only at label-bearing actions, rendered as role-matched speech; each frozen set contains 900 Chinese calls (600 fraud, 300 near-domain non-fraud), controlled text experiments show three classifiers reach perfect Macro-F1 against unrelated or ordinary negatives but drop to 0.65-0.68 with near-domain sibling negatives, and full-set audio and ASR+LLM evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity.
The work presents TeleAntiFraud 2.0, a refreshable Chinese call-audio benchmark organized as monthly frozen snapshots, built with a Mixed-Tree Anti-Fraud Generation Pipeline that turns online fraud-case abstracts into profile-grounded scenarios and expands them into fraud and near-domain lawful sibling dialogues sharing context and diverging only at label-bearing actions, rendered as role-matched speech; each frozen set contains 900 Chinese calls (600 fraud, 300 near-domain non-fraud), controlled text experiments show three classifiers reach perfect Macro-F1 against unrelated or ordinary negatives but drop to 0.65-0.68 with near-domain sibling negatives, and full-set audio and ASR+LLM evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity.
arXiv The work presents a multi-step framework that first corrects word segmentation with ChatGPT-4o, then fills missing entities via context-free lexicons with a minimum frequency of three, and finally applies frequency-based iterative self-training with a dual threshold on logits probabilities and the 90th percentile of self-attention scores, using a multilingual XLM-RoBERTa-large model to select candidates by F1; on manually revised validation/test splits for Urdu MK-PUCIT, Shahmukhi (Western Punjabi), and Sindhi SiNER, fine-tuning XLM-RoBERTa-large yields test micro-F1 gains of 3.96, 1.40, and 1.44 points, while ChatGPT-4o zero/few-shot NER remains below the supervised model.
The work presents a multi-step framework that first corrects word segmentation with ChatGPT-4o, then fills missing entities via context-free lexicons with a minimum frequency of three, and finally applies frequency-based iterative self-training with a dual threshold on logits probabilities and the 90th percentile of self-attention scores, using a multilingual XLM-RoBERTa-large model to select candidates by F1; on manually revised validation/test splits for Urdu MK-PUCIT, Shahmukhi (Western Punjabi), and Sindhi SiNER, fine-tuning XLM-RoBERTa-large yields test micro-F1 gains of 3.96, 1.40, and 1.44 points, while ChatGPT-4o zero/few-shot NER remains below the supervised model.
The work presents a multi-step framework that first corrects word segmentation with ChatGPT-4o, then fills missing entities via context-free lexicons with a minimum frequency of three, and finally applies frequency-based iterative self-training with a dual threshold on logits probabilities and the 90th percentile of self-attention scores, using a multilingual XLM-RoBERTa-large model to select candidates by F1; on manually revised validation/test splits for Urdu MK-PUCIT, Shahmukhi (Western Punjabi), and Sindhi SiNER, fine-tuning XLM-RoBERTa-large yields test micro-F1 gains of 3.96, 1.40, and 1.44 points, while ChatGPT-4o zero/few-shot NER remains below the supervised model.
The work presents a multi-step framework that first corrects word segmentation with ChatGPT-4o, then fills missing entities via context-free lexicons with a minimum frequency of three, and finally applies frequency-based iterative self-training with a dual threshold on logits probabilities and the 90th percentile of self-attention scores, using a multilingual XLM-RoBERTa-large model to select candidates by F1; on manually revised validation/test splits for Urdu MK-PUCIT, Shahmukhi (Western Punjabi), and Sindhi SiNER, fine-tuning XLM-RoBERTa-large yields test micro-F1 gains of 3.96, 1.40, and 1.44 points, while ChatGPT-4o zero/few-shot NER remains below the supervised model.
arXiv This work adapts the board game Clue into a text-based multi-agent environment with six LLM players (three each from GPT-4o-mini and Gemini-2.5-Flash) across 18 games and introduces an external possibility-matrix tool (YES/NO/MAYBE cells plus an accusation_ready signal) that externalizes belief-state tracking, yielding near-perfect accusation accuracy for tool-augmented agents (Gemini-2.5-Flash 1.00, GPT-4o-mini 0.96), raising non-tool agents in mixed games from about 0.81/3 to about 2.71/3, while leaving agents' autonomy over accusation timing unchanged.
This work adapts the board game Clue into a text-based multi-agent environment with six LLM players (three each from GPT-4o-mini and Gemini-2.5-Flash) across 18 games and introduces an external possibility-matrix tool (YES/NO/MAYBE cells plus an accusation_ready signal) that externalizes belief-state tracking, yielding near-perfect accusation accuracy for tool-augmented agents (Gemini-2.5-Flash 1.00, GPT-4o-mini 0.96), raising non-tool agents in mixed games from about 0.81/3 to about 2.71/3, while leaving agents' autonomy over accusation timing unchanged.
This work adapts the board game Clue into a text-based multi-agent environment with six LLM players (three each from GPT-4o-mini and Gemini-2.5-Flash) across 18 games and introduces an external possibility-matrix tool (YES/NO/MAYBE cells plus an accusation_ready signal) that externalizes belief-state tracking, yielding near-perfect accusation accuracy for tool-augmented agents (Gemini-2.5-Flash 1.00, GPT-4o-mini 0.96), raising non-tool agents in mixed games from about 0.81/3 to about 2.71/3, while leaving agents' autonomy over accusation timing unchanged.
This work adapts the board game Clue into a text-based multi-agent environment with six LLM players (three each from GPT-4o-mini and Gemini-2.5-Flash) across 18 games and introduces an external possibility-matrix tool (YES/NO/MAYBE cells plus an accusation_ready signal) that externalizes belief-state tracking, yielding near-perfect accusation accuracy for tool-augmented agents (Gemini-2.5-Flash 1.00, GPT-4o-mini 0.96), raising non-tool agents in mixed games from about 0.81/3 to about 2.71/3, while leaving agents' autonomy over accusation timing unchanged.
arXiv This paper designs an iteratively refined prompt template that lets six contemporary large language models translate informal video-game testing goals written in natural language into PDDL goals for classical planning, and systematically evaluates them on 90 natural language goals (45 expressible and 45 inexpressible across eight domains) for correctness, speed, and error tendencies, finding that all models exceed 92% correctness, with Gemini 2.5 Flash highest at 96% and fewest false positives, while GPT-4.1 is fastest.
This paper designs an iteratively refined prompt template that lets six contemporary large language models translate informal video-game testing goals written in natural language into PDDL goals for classical planning, and systematically evaluates them on 90 natural language goals (45 expressible and 45 inexpressible across eight domains) for correctness, speed, and error tendencies, finding that all models exceed 92% correctness, with Gemini 2.5 Flash highest at 96% and fewest false positives, while GPT-4.1 is fastest.
This paper designs an iteratively refined prompt template that lets six contemporary large language models translate informal video-game testing goals written in natural language into PDDL goals for classical planning, and systematically evaluates them on 90 natural language goals (45 expressible and 45 inexpressible across eight domains) for correctness, speed, and error tendencies, finding that all models exceed 92% correctness, with Gemini 2.5 Flash highest at 96% and fewest false positives, while GPT-4.1 is fastest.
This paper designs an iteratively refined prompt template that lets six contemporary large language models translate informal video-game testing goals written in natural language into PDDL goals for classical planning, and systematically evaluates them on 90 natural language goals (45 expressible and 45 inexpressible across eight domains) for correctness, speed, and error tendencies, finding that all models exceed 92% correctness, with Gemini 2.5 Flash highest at 96% and fewest false positives, while GPT-4.1 is fastest.
NVIDIA Research Emerald AI, Google and NVIDIA announced the launch of the AI Energy Management Alliance (AEMA), a coalition convening the full AI and power value chain around a technology-neutral, performance-based approach that lets data centers dynamically manage electricity use to speed interconnection, strengthen reliability and protect affordability.
Emerald AI, Google and NVIDIA announced the launch of the AI Energy Management Alliance (AEMA), a coalition convening the full AI and power value chain around a technology-neutral, performance-based approach that lets data centers dynamically manage electricity use to speed interconnection, strengthen reliability and protect affordability.
Emerald AI, Google and NVIDIA announced the launch of the AI Energy Management Alliance (AEMA), a coalition convening the full AI and power value chain around a technology-neutral, performance-based approach that lets data centers dynamically manage electricity use to speed interconnection, strengthen reliability and protect affordability.
Emerald AI, Google and NVIDIA announced the launch of the AI Energy Management Alliance (AEMA), a coalition convening the full AI and power value chain around a technology-neutral, performance-based approach that lets data centers dynamically manage electricity use to speed interconnection, strengthen reliability and protect affordability.
OpenAI OpenAI published an article introducing its AI-powered advertising experiences, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify, aiming to explore new forms of advertising in the AI era.
OpenAI published an article introducing its AI-powered advertising experiences, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify, aiming to explore new forms of advertising in the AI era.
OpenAI published an article introducing its AI-powered advertising experiences, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify, aiming to explore new forms of advertising in the AI era.
OpenAI published an article introducing its AI-powered advertising experiences, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify, aiming to explore new forms of advertising in the AI era.
MIT Technology Review This MIT Technology Review Insights conversation produced in partnership with Syensqo records the views of Mike Finelli, Syensqo's chief technology and innovation officer and chief North America officer: AI is pushing semiconductors and data centers toward physical limits, which piles up more simultaneous requirements on advanced materials, and Syensqo is responding by developing materials for high-voltage data center architectures, advanced sealing materials for semiconductor manufacturing, and thermal-management solutions including direct immersion cooling fluids, while working with Microsoft on AI agents that digitally synthesize millions of candidate molecules, predict their performance through physics-based simulation, and rank them down to roughly a hundred candidates for laboratory
This MIT Technology Review Insights conversation produced in partnership with Syensqo records the views of Mike Finelli, Syensqo's chief technology and innovation officer and chief North America officer: AI is pushing semiconductors and data centers toward physical limits, which piles up more simultaneous requirements on advanced materials, and Syensqo is responding by developing materials for high-voltage data center architectures, advanced sealing materials for semiconductor manufacturing, and thermal-management solutions including direct immersion cooling fluids, while working with Microsoft on AI agents that digitally synthesize millions of candidate molecules, predict their performance through physics-based simulation, and rank them down to roughly a hundred candidates for laboratory
This MIT Technology Review Insights conversation produced in partnership with Syensqo records the views of Mike Finelli, Syensqo's chief technology and innovation officer and chief North America officer: AI is pushing semiconductors and data centers toward physical limits, which piles up more simultaneous requirements on advanced materials, and Syensqo is responding by developing materials for high-voltage data center architectures, advanced sealing materials for semiconductor manufacturing, and thermal-management solutions including direct immersion cooling fluids, while working with Microsoft on AI agents that digitally synthesize millions of candidate molecules, predict their performance through physics-based simulation, and rank them down to roughly a hundred candidates for laboratory
This MIT Technology Review Insights conversation produced in partnership with Syensqo records the views of Mike Finelli, Syensqo's chief technology and innovation officer and chief North America officer: AI is pushing semiconductors and data centers toward physical limits, which piles up more simultaneous requirements on advanced materials, and Syensqo is responding by developing materials for high-voltage data center architectures, advanced sealing materials for semiconductor manufacturing, and thermal-management solutions including direct immersion cooling fluids, while working with Microsoft on AI agents that digitally synthesize millions of candidate molecules, predict their performance through physics-based simulation, and rank them down to roughly a hundred candidates for laboratory
MIT Technology Review This is an edition of MIT Technology Review's daily newsletter The Download, rounding up technology news, with highlights including a financial analysis that starts from hyperscalers' nearly $1.1 trillion in data-center spending through 2027 and asks how fast their earnings must grow to break even by 2030, and the OpenAI Foundation's announcement that it will fund policy analyst Ruxandra Teslo's idea of obtaining regulatory filings and safety data from failed biotech companies through bankruptcy proceedings to build what she calls "biotech's lost archive."
This is an edition of MIT Technology Review's daily newsletter The Download, rounding up technology news, with highlights including a financial analysis that starts from hyperscalers' nearly $1.1 trillion in data-center spending through 2027 and asks how fast their earnings must grow to break even by 2030, and the OpenAI Foundation's announcement that it will fund policy analyst Ruxandra Teslo's idea of obtaining regulatory filings and safety data from failed biotech companies through bankruptcy proceedings to build what she calls "biotech's lost archive."
This is an edition of MIT Technology Review's daily newsletter The Download, rounding up technology news, with highlights including a financial analysis that starts from hyperscalers' nearly $1.1 trillion in data-center spending through 2027 and asks how fast their earnings must grow to break even by 2030, and the OpenAI Foundation's announcement that it will fund policy analyst Ruxandra Teslo's idea of obtaining regulatory filings and safety data from failed biotech companies through bankruptcy proceedings to build what she calls "biotech's lost archive."
This is an edition of MIT Technology Review's daily newsletter The Download, rounding up technology news, with highlights including a financial analysis that starts from hyperscalers' nearly $1.1 trillion in data-center spending through 2027 and asks how fast their earnings must grow to break even by 2030, and the OpenAI Foundation's announcement that it will fund policy analyst Ruxandra Teslo's idea of obtaining regulatory filings and safety data from failed biotech companies through bankruptcy proceedings to build what she calls "biotech's lost archive."
OpenAI This OpenAI article explains how ChatGPT Work and Codex analytics help teams understand AI usage and spend, identify training needs, and connect adoption to business outcomes.
This OpenAI article explains how ChatGPT Work and Codex analytics help teams understand AI usage and spend, identify training needs, and connect adoption to business outcomes.
This OpenAI article explains how ChatGPT Work and Codex analytics help teams understand AI usage and spend, identify training needs, and connect adoption to business outcomes.
This OpenAI article explains how ChatGPT Work and Codex analytics help teams understand AI usage and spend, identify training needs, and connect adoption to business outcomes.
Mistral AI Mistral and Mozilla announced a partnership in which Mistral models now power Firefox's AI browsing assistant Smart Window (beta), starting with users in France and North America and expected to reach the United Kingdom and Germany later this year, while emphasizing open distribution, fine-tuning on regional languages and dialects, conversations not saved on Mozilla's servers by default with zero data retention, and bringing sovereign AI to everyday users.
Mistral and Mozilla announced a partnership in which Mistral models now power Firefox's AI browsing assistant Smart Window (beta), starting with users in France and North America and expected to reach the United Kingdom and Germany later this year, while emphasizing open distribution, fine-tuning on regional languages and dialects, conversations not saved on Mozilla's servers by default with zero data retention, and bringing sovereign AI to everyday users.
Mistral and Mozilla announced a partnership in which Mistral models now power Firefox's AI browsing assistant Smart Window (beta), starting with users in France and North America and expected to reach the United Kingdom and Germany later this year, while emphasizing open distribution, fine-tuning on regional languages and dialects, conversations not saved on Mozilla's servers by default with zero data retention, and bringing sovereign AI to everyday users.
Mistral and Mozilla announced a partnership in which Mistral models now power Firefox's AI browsing assistant Smart Window (beta), starting with users in France and North America and expected to reach the United Kingdom and Germany later this year, while emphasizing open distribution, fine-tuning on regional languages and dialects, conversations not saved on Mozilla's servers by default with zero data retention, and bringing sovereign AI to everyday users.
OpenAI An analysis from New OpenAI Economic Research indicates that workers are using AI for purposes beyond their traditional job roles, and identifies which of these new activities become recurring parts of their work.
An analysis from New OpenAI Economic Research indicates that workers are using AI for purposes beyond their traditional job roles, and identifies which of these new activities become recurring parts of their work.
An analysis from New OpenAI Economic Research indicates that workers are using AI for purposes beyond their traditional job roles, and identifies which of these new activities become recurring parts of their work.
An analysis from New OpenAI Economic Research indicates that workers are using AI for purposes beyond their traditional job roles, and identifies which of these new activities become recurring parts of their work.
NVIDIA Research Manchester physicists David Topping and Hao Zhang, working with the NVIDIA Earth-2 team, moved generative models originally built for weather into air quality forecasting: they trained the Earth-2 CorrDiff generative downscaling model on a year of hourly UK chemistry-climate simulation data, completing training in two days on a single eight-GPU node of Isambard-AI to produce a UK-wide pollution model at 2-3 square kilometer resolution, added Earth-2 StormCast for time-dependent forecasts that directly use air quality observations, demonstrated inference and smaller training runs on the DGX Spark desktop AI system, and plan to release open source training data and workflows so other countries and regions can train their own models.
Manchester physicists David Topping and Hao Zhang, working with the NVIDIA Earth-2 team, moved generative models originally built for weather into air quality forecasting: they trained the Earth-2 CorrDiff generative downscaling model on a year of hourly UK chemistry-climate simulation data, completing training in two days on a single eight-GPU node of Isambard-AI to produce a UK-wide pollution model at 2-3 square kilometer resolution, added Earth-2 StormCast for time-dependent forecasts that directly use air quality observations, demonstrated inference and smaller training runs on the DGX Spark desktop AI system, and plan to release open source training data and workflows so other countries and regions can train their own models.
Manchester physicists David Topping and Hao Zhang, working with the NVIDIA Earth-2 team, moved generative models originally built for weather into air quality forecasting: they trained the Earth-2 CorrDiff generative downscaling model on a year of hourly UK chemistry-climate simulation data, completing training in two days on a single eight-GPU node of Isambard-AI to produce a UK-wide pollution model at 2-3 square kilometer resolution, added Earth-2 StormCast for time-dependent forecasts that directly use air quality observations, demonstrated inference and smaller training runs on the DGX Spark desktop AI system, and plan to release open source training data and workflows so other countries and regions can train their own models.
Manchester physicists David Topping and Hao Zhang, working with the NVIDIA Earth-2 team, moved generative models originally built for weather into air quality forecasting: they trained the Earth-2 CorrDiff generative downscaling model on a year of hourly UK chemistry-climate simulation data, completing training in two days on a single eight-GPU node of Isambard-AI to produce a UK-wide pollution model at 2-3 square kilometer resolution, added Earth-2 StormCast for time-dependent forecasts that directly use air quality observations, demonstrated inference and smaller training runs on the DGX Spark desktop AI system, and plan to release open source training data and workflows so other countries and regions can train their own models.
ORBi UMONS This work builds and releases the GHAgentFiles dataset together with the open-source extraction tool cofee, identifying 27,717 repositories with coding agent files out of 165,281 candidates and extracting 142,294 context, skill and subagent file histories (401,873 file revisions across 185,795 commits, spanning 22 November 2022 to 1 July 2026), and it presents a preliminary observational analysis showing that coding agent files emerge around early 2025, peak in additions and modifications in early 2026, and that the generic AGENTS.md has become the most popular naming practice.
This work builds and releases the GHAgentFiles dataset together with the open-source extraction tool cofee, identifying 27,717 repositories with coding agent files out of 165,281 candidates and extracting 142,294 context, skill and subagent file histories (401,873 file revisions across 185,795 commits, spanning 22 November 2022 to 1 July 2026), and it presents a preliminary observational analysis showing that coding agent files emerge around early 2025, peak in additions and modifications in early 2026, and that the generic AGENTS.md has become the most popular naming practice.
This work builds and releases the GHAgentFiles dataset together with the open-source extraction tool cofee, identifying 27,717 repositories with coding agent files out of 165,281 candidates and extracting 142,294 context, skill and subagent file histories (401,873 file revisions across 185,795 commits, spanning 22 November 2022 to 1 July 2026), and it presents a preliminary observational analysis showing that coding agent files emerge around early 2025, peak in additions and modifications in early 2026, and that the generic AGENTS.md has become the most popular naming practice.
This work builds and releases the GHAgentFiles dataset together with the open-source extraction tool cofee, identifying 27,717 repositories with coding agent files out of 165,281 candidates and extracting 142,294 context, skill and subagent file histories (401,873 file revisions across 185,795 commits, spanning 22 November 2022 to 1 July 2026), and it presents a preliminary observational analysis showing that coding agent files emerge around early 2025, peak in additions and modifications in early 2026, and that the generic AGENTS.md has become the most popular naming practice.
Diagnosis (Berlin, Germany) In Fall 2024, 175 second-year medical students completed three MAESSCR multi-agent LLM clinical encounters as coursework, and six clinician-educators rated 120 randomly sampled transcripts with a dichotomous tool, finding 92% (110/120) adherence to scripted details, 2.5% (3/120) diagnosis-changing information, 12% (14/120) unrealistic patient portrayal, and 17% (20/120) technical issues, while the most prominent problem was agents interpreting findings before students could, occurring in 28% (34/120) of encounters for history/physical exam agents and 59% (71/120) for diagnostics/management agents, with most disruptions judged minor.
In Fall 2024, 175 second-year medical students completed three MAESSCR multi-agent LLM clinical encounters as coursework, and six clinician-educators rated 120 randomly sampled transcripts with a dichotomous tool, finding 92% (110/120) adherence to scripted details, 2.5% (3/120) diagnosis-changing information, 12% (14/120) unrealistic patient portrayal, and 17% (20/120) technical issues, while the most prominent problem was agents interpreting findings before students could, occurring in 28% (34/120) of encounters for history/physical exam agents and 59% (71/120) for diagnostics/management agents, with most disruptions judged minor.
In Fall 2024, 175 second-year medical students completed three MAESSCR multi-agent LLM clinical encounters as coursework, and six clinician-educators rated 120 randomly sampled transcripts with a dichotomous tool, finding 92% (110/120) adherence to scripted details, 2.5% (3/120) diagnosis-changing information, 12% (14/120) unrealistic patient portrayal, and 17% (20/120) technical issues, while the most prominent problem was agents interpreting findings before students could, occurring in 28% (34/120) of encounters for history/physical exam agents and 59% (71/120) for diagnostics/management agents, with most disruptions judged minor.
In Fall 2024, 175 second-year medical students completed three MAESSCR multi-agent LLM clinical encounters as coursework, and six clinician-educators rated 120 randomly sampled transcripts with a dichotomous tool, finding 92% (110/120) adherence to scripted details, 2.5% (3/120) diagnosis-changing information, 12% (14/120) unrealistic patient portrayal, and 17% (20/120) technical issues, while the most prominent problem was agents interpreting findings before students could, occurring in 28% (34/120) of encounters for history/physical exam agents and 59% (71/120) for diagnostics/management agents, with most disruptions judged minor.
Journal of medical imaging and radiation sciences This narrative review searched PubMed, Scopus and Web of Science (2012-2024) to map the role of radiomics and artificial intelligence in neuro-oncology across diagnostic, prognostic and therapeutic potential, reporting that radiomics can automatically extract quantitative features from MRI, CT and PET/CT that are not visible to the human eye, has been shown useful for pre-surgical classification of gliomas, survival prediction and differentiation between tumour progression and pseudo-progression, may improve diagnostic accuracy and support patient stratification by molecular biomarkers such as IDH and MGMT when integrated with deep learning, and links image phenotypes to genetic alterations through radiogenomics to enhance personalised medicine, while limitations persist from lack of proto
This narrative review searched PubMed, Scopus and Web of Science (2012-2024) to map the role of radiomics and artificial intelligence in neuro-oncology across diagnostic, prognostic and therapeutic potential, reporting that radiomics can automatically extract quantitative features from MRI, CT and PET/CT that are not visible to the human eye, has been shown useful for pre-surgical classification of gliomas, survival prediction and differentiation between tumour progression and pseudo-progression, may improve diagnostic accuracy and support patient stratification by molecular biomarkers such as IDH and MGMT when integrated with deep learning, and links image phenotypes to genetic alterations through radiogenomics to enhance personalised medicine, while limitations persist from lack of proto
This narrative review searched PubMed, Scopus and Web of Science (2012-2024) to map the role of radiomics and artificial intelligence in neuro-oncology across diagnostic, prognostic and therapeutic potential, reporting that radiomics can automatically extract quantitative features from MRI, CT and PET/CT that are not visible to the human eye, has been shown useful for pre-surgical classification of gliomas, survival prediction and differentiation between tumour progression and pseudo-progression, may improve diagnostic accuracy and support patient stratification by molecular biomarkers such as IDH and MGMT when integrated with deep learning, and links image phenotypes to genetic alterations through radiogenomics to enhance personalised medicine, while limitations persist from lack of proto
This narrative review searched PubMed, Scopus and Web of Science (2012-2024) to map the role of radiomics and artificial intelligence in neuro-oncology across diagnostic, prognostic and therapeutic potential, reporting that radiomics can automatically extract quantitative features from MRI, CT and PET/CT that are not visible to the human eye, has been shown useful for pre-surgical classification of gliomas, survival prediction and differentiation between tumour progression and pseudo-progression, may improve diagnostic accuracy and support patient stratification by molecular biomarkers such as IDH and MGMT when integrated with deep learning, and links image phenotypes to genetic alterations through radiogenomics to enhance personalised medicine, while limitations persist from lack of proto
Science advances This work fabricates a chip-scale microwave diffractive neural network (MDNN) in a GaAs semiconductor process, integrating cascaded couplers and phase shifters to implement a diffraction network within a millimeter-scale footprint, reducing the size of conventional MDNNs by over four orders of magnitude, achieving a computational latency of 2.05 ns and a system-level energy efficiency of 0.83 TOPS/W, and reaching more than 86% accuracy across three functional prototypes—MNIST handwritten digit recognition, multi-user interference suppression, and real-time obstacle perception for drones—thereby validating the chip's capability to directly perform both digital image processing and in-situ electromagnetic information processing in the microwave domain.
This work fabricates a chip-scale microwave diffractive neural network (MDNN) in a GaAs semiconductor process, integrating cascaded couplers and phase shifters to implement a diffraction network within a millimeter-scale footprint, reducing the size of conventional MDNNs by over four orders of magnitude, achieving a computational latency of 2.05 ns and a system-level energy efficiency of 0.83 TOPS/W, and reaching more than 86% accuracy across three functional prototypes—MNIST handwritten digit recognition, multi-user interference suppression, and real-time obstacle perception for drones—thereby validating the chip's capability to directly perform both digital image processing and in-situ electromagnetic information processing in the microwave domain.
This work fabricates a chip-scale microwave diffractive neural network (MDNN) in a GaAs semiconductor process, integrating cascaded couplers and phase shifters to implement a diffraction network within a millimeter-scale footprint, reducing the size of conventional MDNNs by over four orders of magnitude, achieving a computational latency of 2.05 ns and a system-level energy efficiency of 0.83 TOPS/W, and reaching more than 86% accuracy across three functional prototypes—MNIST handwritten digit recognition, multi-user interference suppression, and real-time obstacle perception for drones—thereby validating the chip's capability to directly perform both digital image processing and in-situ electromagnetic information processing in the microwave domain.
This work fabricates a chip-scale microwave diffractive neural network (MDNN) in a GaAs semiconductor process, integrating cascaded couplers and phase shifters to implement a diffraction network within a millimeter-scale footprint, reducing the size of conventional MDNNs by over four orders of magnitude, achieving a computational latency of 2.05 ns and a system-level energy efficiency of 0.83 TOPS/W, and reaching more than 86% accuracy across three functional prototypes—MNIST handwritten digit recognition, multi-user interference suppression, and real-time obstacle perception for drones—thereby validating the chip's capability to directly perform both digital image processing and in-situ electromagnetic information processing in the microwave domain.
JMIR Medical Education This article introduces the AWARE (AI use, why, attachment, reality and risk, and effect on functioning) framework as an educational tool to help mental health professionals systematically assess patients' use of conversational AI through five clinically relevant domains, and discusses incorporating it into undergraduate, postgraduate, and continuing professional education along with priorities for future research.
This article introduces the AWARE (AI use, why, attachment, reality and risk, and effect on functioning) framework as an educational tool to help mental health professionals systematically assess patients' use of conversational AI through five clinically relevant domains, and discusses incorporating it into undergraduate, postgraduate, and continuing professional education along with priorities for future research.
This article introduces the AWARE (AI use, why, attachment, reality and risk, and effect on functioning) framework as an educational tool to help mental health professionals systematically assess patients' use of conversational AI through five clinically relevant domains, and discusses incorporating it into undergraduate, postgraduate, and continuing professional education along with priorities for future research.
This article introduces the AWARE (AI use, why, attachment, reality and risk, and effect on functioning) framework as an educational tool to help mental health professionals systematically assess patients' use of conversational AI through five clinically relevant domains, and discusses incorporating it into undergraduate, postgraduate, and continuing professional education along with priorities for future research.
medRxiv Using a Nominal Group Technique session (May 2025, n=6) and a two-round modified Delphi process (April–May 2026, expert panel n=13 across 12 academic medical centers, with 77% Round 1 and 100% Round 2 response rates), this study developed the first specialty-specific AI competency framework for a United States medical specialty, comprising 5 themes (Communicating about AI, Understanding appropriate use cases, Interacting with AI, AI risk management, Cognitive impacts of AI), 5 derived competencies (one-to-one theme-to-competency mapping endorsed by 12 of 13 panelists), 19 subthemes, and 10 retained clinical scenarios, with all themes, derived competencies, and individually rated subthemes meeting pre-specified consensus thresholds and 9 of 10 scenarios reaching consensus while 1 was retain
Using a Nominal Group Technique session (May 2025, n=6) and a two-round modified Delphi process (April–May 2026, expert panel n=13 across 12 academic medical centers, with 77% Round 1 and 100% Round 2 response rates), this study developed the first specialty-specific AI competency framework for a United States medical specialty, comprising 5 themes (Communicating about AI, Understanding appropriate use cases, Interacting with AI, AI risk management, Cognitive impacts of AI), 5 derived competencies (one-to-one theme-to-competency mapping endorsed by 12 of 13 panelists), 19 subthemes, and 10 retained clinical scenarios, with all themes, derived competencies, and individually rated subthemes meeting pre-specified consensus thresholds and 9 of 10 scenarios reaching consensus while 1 was retain
Using a Nominal Group Technique session (May 2025, n=6) and a two-round modified Delphi process (April–May 2026, expert panel n=13 across 12 academic medical centers, with 77% Round 1 and 100% Round 2 response rates), this study developed the first specialty-specific AI competency framework for a United States medical specialty, comprising 5 themes (Communicating about AI, Understanding appropriate use cases, Interacting with AI, AI risk management, Cognitive impacts of AI), 5 derived competencies (one-to-one theme-to-competency mapping endorsed by 12 of 13 panelists), 19 subthemes, and 10 retained clinical scenarios, with all themes, derived competencies, and individually rated subthemes meeting pre-specified consensus thresholds and 9 of 10 scenarios reaching consensus while 1 was retain
Using a Nominal Group Technique session (May 2025, n=6) and a two-round modified Delphi process (April–May 2026, expert panel n=13 across 12 academic medical centers, with 77% Round 1 and 100% Round 2 response rates), this study developed the first specialty-specific AI competency framework for a United States medical specialty, comprising 5 themes (Communicating about AI, Understanding appropriate use cases, Interacting with AI, AI risk management, Cognitive impacts of AI), 5 derived competencies (one-to-one theme-to-competency mapping endorsed by 12 of 13 panelists), 19 subthemes, and 10 retained clinical scenarios, with all themes, derived competencies, and individually rated subthemes meeting pre-specified consensus thresholds and 9 of 10 scenarios reaching consensus while 1 was retain
Clinical therapeutics This narrative review, based on structured searches of PubMed and ScienceDirect (2016-2026, plus a clinical-research-focused search for 2013-2026) supplemented by FDA, EMA, Cochrane Library, and industry sources, examines medical device safety across the full lifecycle from early design and clinical investigation through regulatory review, postmarket surveillance, software change management, cybersecurity, and AI/ML-enabled device oversight, arguing that device safety requires a broader socio-technical framework than drug safety and that safety evidence is shifting from retrospective, passive reporting toward proactive, real-time, lifecycle-based generation.
This narrative review, based on structured searches of PubMed and ScienceDirect (2016-2026, plus a clinical-research-focused search for 2013-2026) supplemented by FDA, EMA, Cochrane Library, and industry sources, examines medical device safety across the full lifecycle from early design and clinical investigation through regulatory review, postmarket surveillance, software change management, cybersecurity, and AI/ML-enabled device oversight, arguing that device safety requires a broader socio-technical framework than drug safety and that safety evidence is shifting from retrospective, passive reporting toward proactive, real-time, lifecycle-based generation.
This narrative review, based on structured searches of PubMed and ScienceDirect (2016-2026, plus a clinical-research-focused search for 2013-2026) supplemented by FDA, EMA, Cochrane Library, and industry sources, examines medical device safety across the full lifecycle from early design and clinical investigation through regulatory review, postmarket surveillance, software change management, cybersecurity, and AI/ML-enabled device oversight, arguing that device safety requires a broader socio-technical framework than drug safety and that safety evidence is shifting from retrospective, passive reporting toward proactive, real-time, lifecycle-based generation.
This narrative review, based on structured searches of PubMed and ScienceDirect (2016-2026, plus a clinical-research-focused search for 2013-2026) supplemented by FDA, EMA, Cochrane Library, and industry sources, examines medical device safety across the full lifecycle from early design and clinical investigation through regulatory review, postmarket surveillance, software change management, cybersecurity, and AI/ML-enabled device oversight, arguing that device safety requires a broader socio-technical framework than drug safety and that safety evidence is shifting from retrospective, passive reporting toward proactive, real-time, lifecycle-based generation.
JMIR Medical Informatics This systematic review searched PubMed, MEDLINE, Embase, PsycINFO, and Web of Science from inception to November 11, 2025, and included 29 studies that developed, validated, or evaluated multivariable prediction models using routinely collected electronic health record or administrative data to predict acute mental status deterioration during adult hospital admissions, all operationalized as delirium; using CHARMS and TRIPOD/TRIPOD-AI for data extraction and PROBAST for risk of bias, it found that the evidence clustered into four overlapping prediction tasks (admission or early-stay risk stratification, perioperative or postoperative prediction, dynamic intensive care unit prediction, and external validation or workflow evaluation of existing tools), that most studies were retrospective co
This systematic review searched PubMed, MEDLINE, Embase, PsycINFO, and Web of Science from inception to November 11, 2025, and included 29 studies that developed, validated, or evaluated multivariable prediction models using routinely collected electronic health record or administrative data to predict acute mental status deterioration during adult hospital admissions, all operationalized as delirium; using CHARMS and TRIPOD/TRIPOD-AI for data extraction and PROBAST for risk of bias, it found that the evidence clustered into four overlapping prediction tasks (admission or early-stay risk stratification, perioperative or postoperative prediction, dynamic intensive care unit prediction, and external validation or workflow evaluation of existing tools), that most studies were retrospective co
This systematic review searched PubMed, MEDLINE, Embase, PsycINFO, and Web of Science from inception to November 11, 2025, and included 29 studies that developed, validated, or evaluated multivariable prediction models using routinely collected electronic health record or administrative data to predict acute mental status deterioration during adult hospital admissions, all operationalized as delirium; using CHARMS and TRIPOD/TRIPOD-AI for data extraction and PROBAST for risk of bias, it found that the evidence clustered into four overlapping prediction tasks (admission or early-stay risk stratification, perioperative or postoperative prediction, dynamic intensive care unit prediction, and external validation or workflow evaluation of existing tools), that most studies were retrospective co
This systematic review searched PubMed, MEDLINE, Embase, PsycINFO, and Web of Science from inception to November 11, 2025, and included 29 studies that developed, validated, or evaluated multivariable prediction models using routinely collected electronic health record or administrative data to predict acute mental status deterioration during adult hospital admissions, all operationalized as delirium; using CHARMS and TRIPOD/TRIPOD-AI for data extraction and PROBAST for risk of bias, it found that the evidence clustered into four overlapping prediction tasks (admission or early-stay risk stratification, perioperative or postoperative prediction, dynamic intensive care unit prediction, and external validation or workflow evaluation of existing tools), that most studies were retrospective co
ACM Computing Surveys This survey organizes the landscape of mixed-precision quantization frameworks for language models (MXPLMs): it first reviews quantization fundamentals including uniform and non-uniform quantizers, quantization granularity, and widely used post-training quantization methods, then categorizes and compares recent MXPLM frameworks by their bit allocation strategies and precision configurations across weights, activations, and key-value caches, contrasts them with earlier mixed-precision methods for deep neural networks to identify strategies that transfer and those that face challenges in the LM setting, and closes with open issues such as hardware-aware design, activation quantization, and scalable optimization for billion-parameter models.
This survey organizes the landscape of mixed-precision quantization frameworks for language models (MXPLMs): it first reviews quantization fundamentals including uniform and non-uniform quantizers, quantization granularity, and widely used post-training quantization methods, then categorizes and compares recent MXPLM frameworks by their bit allocation strategies and precision configurations across weights, activations, and key-value caches, contrasts them with earlier mixed-precision methods for deep neural networks to identify strategies that transfer and those that face challenges in the LM setting, and closes with open issues such as hardware-aware design, activation quantization, and scalable optimization for billion-parameter models.
This survey organizes the landscape of mixed-precision quantization frameworks for language models (MXPLMs): it first reviews quantization fundamentals including uniform and non-uniform quantizers, quantization granularity, and widely used post-training quantization methods, then categorizes and compares recent MXPLM frameworks by their bit allocation strategies and precision configurations across weights, activations, and key-value caches, contrasts them with earlier mixed-precision methods for deep neural networks to identify strategies that transfer and those that face challenges in the LM setting, and closes with open issues such as hardware-aware design, activation quantization, and scalable optimization for billion-parameter models.
This survey organizes the landscape of mixed-precision quantization frameworks for language models (MXPLMs): it first reviews quantization fundamentals including uniform and non-uniform quantizers, quantization granularity, and widely used post-training quantization methods, then categorizes and compares recent MXPLM frameworks by their bit allocation strategies and precision configurations across weights, activations, and key-value caches, contrasts them with earlier mixed-precision methods for deep neural networks to identify strategies that transfer and those that face challenges in the LM setting, and closes with open issues such as hardware-aware design, activation quantization, and scalable optimization for billion-parameter models.
发表出处待核验 This work introduces information satisfaction—a query paired with a reader persona—as an axis of summarization evaluation, and through five perturbation tests plus an expert human evaluation finds that traditional metrics such as ROUGE and BERTScore and LLM-as-judge metrics such as Llama-3.3-70B and Prometheus-7B mostly fail basic perturbation checks and agree with reader preferences at near-chance levels, indicating that existing metrics are insufficient measures of how well a summary serves a specific reader's informational needs.
This work introduces information satisfaction—a query paired with a reader persona—as an axis of summarization evaluation, and through five perturbation tests plus an expert human evaluation finds that traditional metrics such as ROUGE and BERTScore and LLM-as-judge metrics such as Llama-3.3-70B and Prometheus-7B mostly fail basic perturbation checks and agree with reader preferences at near-chance levels, indicating that existing metrics are insufficient measures of how well a summary serves a specific reader's informational needs.
This work introduces information satisfaction—a query paired with a reader persona—as an axis of summarization evaluation, and through five perturbation tests plus an expert human evaluation finds that traditional metrics such as ROUGE and BERTScore and LLM-as-judge metrics such as Llama-3.3-70B and Prometheus-7B mostly fail basic perturbation checks and agree with reader preferences at near-chance levels, indicating that existing metrics are insufficient measures of how well a summary serves a specific reader's informational needs.
This work introduces information satisfaction—a query paired with a reader persona—as an axis of summarization evaluation, and through five perturbation tests plus an expert human evaluation finds that traditional metrics such as ROUGE and BERTScore and LLM-as-judge metrics such as Llama-3.3-70B and Prometheus-7B mostly fail basic perturbation checks and agree with reader preferences at near-chance levels, indicating that existing metrics are insufficient measures of how well a summary serves a specific reader's informational needs.
Figshare This replication package supplies full verifiable materials for Bauer (2026), which asks when a liability commitment can still certify a provider's preserved human fallback capability once generative AI makes the output itself uninformative, mapping a posted cap and agreed-damages term into retained exposure through four legal primitives and deriving the message set {0} u [F, C] in which low types pool at zero, intermediate types separate on a schedule anchored at F, and high types may pool at the ceiling, with separation beginning at the bottom type where law removes the zero-exposure region.
This replication package supplies full verifiable materials for Bauer (2026), which asks when a liability commitment can still certify a provider's preserved human fallback capability once generative AI makes the output itself uninformative, mapping a posted cap and agreed-damages term into retained exposure through four legal primitives and deriving the message set {0} u [F, C] in which low types pool at zero, intermediate types separate on a schedule anchored at F, and high types may pool at the ceiling, with separation beginning at the bottom type where law removes the zero-exposure region.
This replication package supplies full verifiable materials for Bauer (2026), which asks when a liability commitment can still certify a provider's preserved human fallback capability once generative AI makes the output itself uninformative, mapping a posted cap and agreed-damages term into retained exposure through four legal primitives and deriving the message set {0} u [F, C] in which low types pool at zero, intermediate types separate on a schedule anchored at F, and high types may pool at the ceiling, with separation beginning at the bottom type where law removes the zero-exposure region.
This replication package supplies full verifiable materials for Bauer (2026), which asks when a liability commitment can still certify a provider's preserved human fallback capability once generative AI makes the output itself uninformative, mapping a posted cap and agreed-damages term into retained exposure through four legal primitives and deriving the message set {0} u [F, C] in which low types pool at zero, intermediate types separate on a schedule anchored at F, and high types may pool at the ceiling, with separation beginning at the bottom type where law removes the zero-exposure region.
Biologia futura This review states that microbial diversity underpins human health, agricultural productivity, ecological balance, and ecosystem functioning, and it surveys the role of the gut microbial community in immune regulation, metabolism, and disease prevention, the contributions of interactions among plants, fungi, bacteria, and other soil microorganisms to carbon sequestration, nutrient cycling, stress resilience, and sustainable agricultural productivity in terrestrial ecosystems, and emerging microbiome-based therapies such as precision probiotics, postbiotics, faecal microbiota transplantation, and personalized microbiome medicine, while noting that multi-omic techniques, synthetic microbial genomes, microbiome engineering, and artificial intelligence enable emerging uses in agriculture, envi
This review states that microbial diversity underpins human health, agricultural productivity, ecological balance, and ecosystem functioning, and it surveys the role of the gut microbial community in immune regulation, metabolism, and disease prevention, the contributions of interactions among plants, fungi, bacteria, and other soil microorganisms to carbon sequestration, nutrient cycling, stress resilience, and sustainable agricultural productivity in terrestrial ecosystems, and emerging microbiome-based therapies such as precision probiotics, postbiotics, faecal microbiota transplantation, and personalized microbiome medicine, while noting that multi-omic techniques, synthetic microbial genomes, microbiome engineering, and artificial intelligence enable emerging uses in agriculture, envi
This review states that microbial diversity underpins human health, agricultural productivity, ecological balance, and ecosystem functioning, and it surveys the role of the gut microbial community in immune regulation, metabolism, and disease prevention, the contributions of interactions among plants, fungi, bacteria, and other soil microorganisms to carbon sequestration, nutrient cycling, stress resilience, and sustainable agricultural productivity in terrestrial ecosystems, and emerging microbiome-based therapies such as precision probiotics, postbiotics, faecal microbiota transplantation, and personalized microbiome medicine, while noting that multi-omic techniques, synthetic microbial genomes, microbiome engineering, and artificial intelligence enable emerging uses in agriculture, envi
This review states that microbial diversity underpins human health, agricultural productivity, ecological balance, and ecosystem functioning, and it surveys the role of the gut microbial community in immune regulation, metabolism, and disease prevention, the contributions of interactions among plants, fungi, bacteria, and other soil microorganisms to carbon sequestration, nutrient cycling, stress resilience, and sustainable agricultural productivity in terrestrial ecosystems, and emerging microbiome-based therapies such as precision probiotics, postbiotics, faecal microbiota transplantation, and personalized microbiome medicine, while noting that multi-omic techniques, synthetic microbial genomes, microbiome engineering, and artificial intelligence enable emerging uses in agriculture, envi
Climatic Change This study introduces a 'generative debunking' framework that integrates climate contrarian claim classification (CARDS) and fallacy detection (FLICC) into an LLM prompting pipeline, so that a model takes a climate myth as input and produces a debunking that follows the fact-myth-fallacy-fact ('truth sandwich') structure; the authors pair three prompting strategies of increasing complexity with GPT-4, Palm2, and Mixtral, and four authors (including one climate misinformation expert) rate 60 debunkings for 20 myths on fact, fallacy, and structure, finding that GPT-4 with a simple prompt and Mixtral with a structured prompt perform relatively well, that fallacy explanations score higher than facts, and that non-expert annotators show poor agreement with the expert on fact quality.
This study introduces a 'generative debunking' framework that integrates climate contrarian claim classification (CARDS) and fallacy detection (FLICC) into an LLM prompting pipeline, so that a model takes a climate myth as input and produces a debunking that follows the fact-myth-fallacy-fact ('truth sandwich') structure; the authors pair three prompting strategies of increasing complexity with GPT-4, Palm2, and Mixtral, and four authors (including one climate misinformation expert) rate 60 debunkings for 20 myths on fact, fallacy, and structure, finding that GPT-4 with a simple prompt and Mixtral with a structured prompt perform relatively well, that fallacy explanations score higher than facts, and that non-expert annotators show poor agreement with the expert on fact quality.
This study introduces a 'generative debunking' framework that integrates climate contrarian claim classification (CARDS) and fallacy detection (FLICC) into an LLM prompting pipeline, so that a model takes a climate myth as input and produces a debunking that follows the fact-myth-fallacy-fact ('truth sandwich') structure; the authors pair three prompting strategies of increasing complexity with GPT-4, Palm2, and Mixtral, and four authors (including one climate misinformation expert) rate 60 debunkings for 20 myths on fact, fallacy, and structure, finding that GPT-4 with a simple prompt and Mixtral with a structured prompt perform relatively well, that fallacy explanations score higher than facts, and that non-expert annotators show poor agreement with the expert on fact quality.
This study introduces a 'generative debunking' framework that integrates climate contrarian claim classification (CARDS) and fallacy detection (FLICC) into an LLM prompting pipeline, so that a model takes a climate myth as input and produces a debunking that follows the fact-myth-fallacy-fact ('truth sandwich') structure; the authors pair three prompting strategies of increasing complexity with GPT-4, Palm2, and Mixtral, and four authors (including one climate misinformation expert) rate 60 debunkings for 20 myths on fact, fallacy, and structure, finding that GPT-4 with a simple prompt and Mixtral with a structured prompt perform relatively well, that fallacy explanations score higher than facts, and that non-expert annotators show poor agreement with the expert on fact quality.
Current oncology reports This review systematically examines current follow-up protocols and recent developments for nasopharyngeal carcinoma (NPC): on one hand it compares and analyzes major follow-up guidelines in terms of follow-up frequency, imaging modalities (MRI and PET), plasma EBV-DNA monitoring, and functional assessments; on the other hand it elaborates on the application prospects and research progress of genomics, radiomics, and artificial intelligence in NPC surveillance, noting that risk-stratified, individualized follow-up strategies such as those based on conditional survival models can enhance the cost-effectiveness of surveillance, while radiomics and artificial intelligence show promise for improving recurrence risk assessment, prognostic stratification, and individualized surveillance.
This review systematically examines current follow-up protocols and recent developments for nasopharyngeal carcinoma (NPC): on one hand it compares and analyzes major follow-up guidelines in terms of follow-up frequency, imaging modalities (MRI and PET), plasma EBV-DNA monitoring, and functional assessments; on the other hand it elaborates on the application prospects and research progress of genomics, radiomics, and artificial intelligence in NPC surveillance, noting that risk-stratified, individualized follow-up strategies such as those based on conditional survival models can enhance the cost-effectiveness of surveillance, while radiomics and artificial intelligence show promise for improving recurrence risk assessment, prognostic stratification, and individualized surveillance.
This review systematically examines current follow-up protocols and recent developments for nasopharyngeal carcinoma (NPC): on one hand it compares and analyzes major follow-up guidelines in terms of follow-up frequency, imaging modalities (MRI and PET), plasma EBV-DNA monitoring, and functional assessments; on the other hand it elaborates on the application prospects and research progress of genomics, radiomics, and artificial intelligence in NPC surveillance, noting that risk-stratified, individualized follow-up strategies such as those based on conditional survival models can enhance the cost-effectiveness of surveillance, while radiomics and artificial intelligence show promise for improving recurrence risk assessment, prognostic stratification, and individualized surveillance.
This review systematically examines current follow-up protocols and recent developments for nasopharyngeal carcinoma (NPC): on one hand it compares and analyzes major follow-up guidelines in terms of follow-up frequency, imaging modalities (MRI and PET), plasma EBV-DNA monitoring, and functional assessments; on the other hand it elaborates on the application prospects and research progress of genomics, radiomics, and artificial intelligence in NPC surveillance, noting that risk-stratified, individualized follow-up strategies such as those based on conditional survival models can enhance the cost-effectiveness of surveillance, while radiomics and artificial intelligence show promise for improving recurrence risk assessment, prognostic stratification, and individualized surveillance.
BMC Geriatrics Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.
Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.
Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.
Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.
Chemical Society reviews This review addresses active lithium loss during the initial cycling and long-term operation of lithium-ion batteries by establishing a classification framework for cathode lithium-rich compensators (LRCs) based on lithium compensation mechanisms, covering binary, ternary, over-lithiated, sacrificial lithium salt, and sustained-release types; it systematically summarizes their working principles, typical charge compensation pathways, and practical performance, discusses key challenges including air instability, gas evolution, high delithiation voltage, and processing issues, highlights mitigation strategies involving nanoscale engineering, surface coating, defect and doping design, and electrolyte optimization, and further proposes a synergistic multimodal lithium-compensation strategy int
This review addresses active lithium loss during the initial cycling and long-term operation of lithium-ion batteries by establishing a classification framework for cathode lithium-rich compensators (LRCs) based on lithium compensation mechanisms, covering binary, ternary, over-lithiated, sacrificial lithium salt, and sustained-release types; it systematically summarizes their working principles, typical charge compensation pathways, and practical performance, discusses key challenges including air instability, gas evolution, high delithiation voltage, and processing issues, highlights mitigation strategies involving nanoscale engineering, surface coating, defect and doping design, and electrolyte optimization, and further proposes a synergistic multimodal lithium-compensation strategy int
This review addresses active lithium loss during the initial cycling and long-term operation of lithium-ion batteries by establishing a classification framework for cathode lithium-rich compensators (LRCs) based on lithium compensation mechanisms, covering binary, ternary, over-lithiated, sacrificial lithium salt, and sustained-release types; it systematically summarizes their working principles, typical charge compensation pathways, and practical performance, discusses key challenges including air instability, gas evolution, high delithiation voltage, and processing issues, highlights mitigation strategies involving nanoscale engineering, surface coating, defect and doping design, and electrolyte optimization, and further proposes a synergistic multimodal lithium-compensation strategy int
This review addresses active lithium loss during the initial cycling and long-term operation of lithium-ion batteries by establishing a classification framework for cathode lithium-rich compensators (LRCs) based on lithium compensation mechanisms, covering binary, ternary, over-lithiated, sacrificial lithium salt, and sustained-release types; it systematically summarizes their working principles, typical charge compensation pathways, and practical performance, discusses key challenges including air instability, gas evolution, high delithiation voltage, and processing issues, highlights mitigation strategies involving nanoscale engineering, surface coating, defect and doping design, and electrolyte optimization, and further proposes a synergistic multimodal lithium-compensation strategy int
JMIR Formative Research This case report describes a woman in her mid-20s with generalized anxiety disorder and major depressive disorder and a history of strong social, academic, and occupational functioning who developed a pattern of functional dependence on ChatGPT, outsourcing routine cognitive and interpersonal tasks such as composing emails, interpreting social interactions, predicting the future, and making decisions, and becoming increasingly uncomfortable completing such tasks independently; the authors frame this as cognitive offloading, reduced confidence in independent judgment, and reinforcement of externalized thinking using the I-PACE model, and suggest that the unlimited accessibility of AI tools may intensify reassurance seeking and worsen tolerance of uncertainty.
This case report describes a woman in her mid-20s with generalized anxiety disorder and major depressive disorder and a history of strong social, academic, and occupational functioning who developed a pattern of functional dependence on ChatGPT, outsourcing routine cognitive and interpersonal tasks such as composing emails, interpreting social interactions, predicting the future, and making decisions, and becoming increasingly uncomfortable completing such tasks independently; the authors frame this as cognitive offloading, reduced confidence in independent judgment, and reinforcement of externalized thinking using the I-PACE model, and suggest that the unlimited accessibility of AI tools may intensify reassurance seeking and worsen tolerance of uncertainty.
This case report describes a woman in her mid-20s with generalized anxiety disorder and major depressive disorder and a history of strong social, academic, and occupational functioning who developed a pattern of functional dependence on ChatGPT, outsourcing routine cognitive and interpersonal tasks such as composing emails, interpreting social interactions, predicting the future, and making decisions, and becoming increasingly uncomfortable completing such tasks independently; the authors frame this as cognitive offloading, reduced confidence in independent judgment, and reinforcement of externalized thinking using the I-PACE model, and suggest that the unlimited accessibility of AI tools may intensify reassurance seeking and worsen tolerance of uncertainty.
This case report describes a woman in her mid-20s with generalized anxiety disorder and major depressive disorder and a history of strong social, academic, and occupational functioning who developed a pattern of functional dependence on ChatGPT, outsourcing routine cognitive and interpersonal tasks such as composing emails, interpreting social interactions, predicting the future, and making decisions, and becoming increasingly uncomfortable completing such tasks independently; the authors frame this as cognitive offloading, reduced confidence in independent judgment, and reinforcement of externalized thinking using the I-PACE model, and suggest that the unlimited accessibility of AI tools may intensify reassurance seeking and worsen tolerance of uncertainty.
Mendeley Data This work assembles a corrosion inhibitor dataset from 5,597 publications using a multi-agent extraction pipeline, releasing a manually verified IE_datasets.xlsx, LLM-extracted json_extracted.zip, and IE_PH_4_materials.xlsx for model construction, with fields covering reference metadata, inhibitor name and composition, anodic/cathodic/mixed type, film mechanism, SMILES, corrosion material name, grade, composition, processing and heat treatment, corrosive medium, medium type, concentration and temperature, and test temperature, test time, inhibitor concentration, test method and inhibition efficiency percentage.
This work assembles a corrosion inhibitor dataset from 5,597 publications using a multi-agent extraction pipeline, releasing a manually verified IE_datasets.xlsx, LLM-extracted json_extracted.zip, and IE_PH_4_materials.xlsx for model construction, with fields covering reference metadata, inhibitor name and composition, anodic/cathodic/mixed type, film mechanism, SMILES, corrosion material name, grade, composition, processing and heat treatment, corrosive medium, medium type, concentration and temperature, and test temperature, test time, inhibitor concentration, test method and inhibition efficiency percentage.
This work assembles a corrosion inhibitor dataset from 5,597 publications using a multi-agent extraction pipeline, releasing a manually verified IE_datasets.xlsx, LLM-extracted json_extracted.zip, and IE_PH_4_materials.xlsx for model construction, with fields covering reference metadata, inhibitor name and composition, anodic/cathodic/mixed type, film mechanism, SMILES, corrosion material name, grade, composition, processing and heat treatment, corrosive medium, medium type, concentration and temperature, and test temperature, test time, inhibitor concentration, test method and inhibition efficiency percentage.
This work assembles a corrosion inhibitor dataset from 5,597 publications using a multi-agent extraction pipeline, releasing a manually verified IE_datasets.xlsx, LLM-extracted json_extracted.zip, and IE_PH_4_materials.xlsx for model construction, with fields covering reference metadata, inhibitor name and composition, anodic/cathodic/mixed type, film mechanism, SMILES, corrosion material name, grade, composition, processing and heat treatment, corrosive medium, medium type, concentration and temperature, and test temperature, test time, inhibitor concentration, test method and inhibition efficiency percentage.
ACM Journal on Computing and Sustainable Societies Using a 2023 workshop with 49 students as a motivating example, this paper critically reflects on the energy costs of using genAI in design education and develops a set of five alternative stances, with related actions, to support the conscious use of genAI in design education.
Using a 2023 workshop with 49 students as a motivating example, this paper critically reflects on the energy costs of using genAI in design education and develops a set of five alternative stances, with related actions, to support the conscious use of genAI in design education.
Using a 2023 workshop with 49 students as a motivating example, this paper critically reflects on the energy costs of using genAI in design education and develops a set of five alternative stances, with related actions, to support the conscious use of genAI in design education.
Using a 2023 workshop with 49 students as a motivating example, this paper critically reflects on the energy costs of using genAI in design education and develops a set of five alternative stances, with related actions, to support the conscious use of genAI in design education.
Bioscience Reports This review systematically examines how DNA methylation, histone modifications, chromatin remodeling, and non-coding RNA networks drive chemoresistance and recurrence in ovarian cancer, summarizes clinical trial results of epigenetic agents (DNMT, HDAC, and EZH2 inhibitors) combined with PARP inhibitors or immunotherapy, and reviews biomarker advances based on circulating cfDNA methylation, circulating miRNAs, and AI-based liquid biopsy platforms, proposing a precision oncology framework for patient stratification and real-time monitoring of chemoresistance.
This review systematically examines how DNA methylation, histone modifications, chromatin remodeling, and non-coding RNA networks drive chemoresistance and recurrence in ovarian cancer, summarizes clinical trial results of epigenetic agents (DNMT, HDAC, and EZH2 inhibitors) combined with PARP inhibitors or immunotherapy, and reviews biomarker advances based on circulating cfDNA methylation, circulating miRNAs, and AI-based liquid biopsy platforms, proposing a precision oncology framework for patient stratification and real-time monitoring of chemoresistance.
This review systematically examines how DNA methylation, histone modifications, chromatin remodeling, and non-coding RNA networks drive chemoresistance and recurrence in ovarian cancer, summarizes clinical trial results of epigenetic agents (DNMT, HDAC, and EZH2 inhibitors) combined with PARP inhibitors or immunotherapy, and reviews biomarker advances based on circulating cfDNA methylation, circulating miRNAs, and AI-based liquid biopsy platforms, proposing a precision oncology framework for patient stratification and real-time monitoring of chemoresistance.
This review systematically examines how DNA methylation, histone modifications, chromatin remodeling, and non-coding RNA networks drive chemoresistance and recurrence in ovarian cancer, summarizes clinical trial results of epigenetic agents (DNMT, HDAC, and EZH2 inhibitors) combined with PARP inhibitors or immunotherapy, and reviews biomarker advances based on circulating cfDNA methylation, circulating miRNAs, and AI-based liquid biopsy platforms, proposing a precision oncology framework for patient stratification and real-time monitoring of chemoresistance.
Neurotherapeutics : the journal of the American Society for Experimental NeuroTherapeutics This review proposes and organizes a framework of emerging clinical trial models for central nervous system tumors, including master protocol designs, Bayesian adaptive frameworks, trials as active discovery platforms embedding longitudinal tissue sampling, window-of-opportunity designs and multi-omic profiling, and decentralized models with artificial intelligence tools, arguing that trials should be reimagined as dynamic, biologically integrated, learning-based systems rather than static tests of individual agents in order to accelerate therapeutic progress in neuro-oncology.
This review proposes and organizes a framework of emerging clinical trial models for central nervous system tumors, including master protocol designs, Bayesian adaptive frameworks, trials as active discovery platforms embedding longitudinal tissue sampling, window-of-opportunity designs and multi-omic profiling, and decentralized models with artificial intelligence tools, arguing that trials should be reimagined as dynamic, biologically integrated, learning-based systems rather than static tests of individual agents in order to accelerate therapeutic progress in neuro-oncology.
This review proposes and organizes a framework of emerging clinical trial models for central nervous system tumors, including master protocol designs, Bayesian adaptive frameworks, trials as active discovery platforms embedding longitudinal tissue sampling, window-of-opportunity designs and multi-omic profiling, and decentralized models with artificial intelligence tools, arguing that trials should be reimagined as dynamic, biologically integrated, learning-based systems rather than static tests of individual agents in order to accelerate therapeutic progress in neuro-oncology.
This review proposes and organizes a framework of emerging clinical trial models for central nervous system tumors, including master protocol designs, Bayesian adaptive frameworks, trials as active discovery platforms embedding longitudinal tissue sampling, window-of-opportunity designs and multi-omic profiling, and decentralized models with artificial intelligence tools, arguing that trials should be reimagined as dynamic, biologically integrated, learning-based systems rather than static tests of individual agents in order to accelerate therapeutic progress in neuro-oncology.
Trends in pharmacological sciences The article notes that while artificial intelligence accelerates drug discovery it also creates regulatory gaps, with more than 100 AI-assisted pipelines in trials and frameworks lacking enforceable standards; it therefore proposes a risk-tiered framework that distinguishes discovery AI from evidence-generating AI, mandates impact assessments and Investigational New Drug disclosure when AI influences decisions, and transforms guidelines into binding, risk-proportionate regulation.
The article notes that while artificial intelligence accelerates drug discovery it also creates regulatory gaps, with more than 100 AI-assisted pipelines in trials and frameworks lacking enforceable standards; it therefore proposes a risk-tiered framework that distinguishes discovery AI from evidence-generating AI, mandates impact assessments and Investigational New Drug disclosure when AI influences decisions, and transforms guidelines into binding, risk-proportionate regulation.
The article notes that while artificial intelligence accelerates drug discovery it also creates regulatory gaps, with more than 100 AI-assisted pipelines in trials and frameworks lacking enforceable standards; it therefore proposes a risk-tiered framework that distinguishes discovery AI from evidence-generating AI, mandates impact assessments and Investigational New Drug disclosure when AI influences decisions, and transforms guidelines into binding, risk-proportionate regulation.
The article notes that while artificial intelligence accelerates drug discovery it also creates regulatory gaps, with more than 100 AI-assisted pipelines in trials and frameworks lacking enforceable standards; it therefore proposes a risk-tiered framework that distinguishes discovery AI from evidence-generating AI, mandates impact assessments and Investigational New Drug disclosure when AI influences decisions, and transforms guidelines into binding, risk-proportionate regulation.
European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - This narrative review, based on a structured search of PubMed/MEDLINE, Scopus, and Web of Science, maps the clinical applications of artificial intelligence across otolaryngology subspecialties, noting that deep learning shows potential in sinonasal disease detection, automated image segmentation, lymph node metastasis prediction, thyroid nodule classification, and prognostic modeling in head and neck cancer, while multimodal systems and generative large language models are emerging in medical education, image interpretation, differential diagnosis, and clinical decision support; however, limited external validation, retrospective designs, dataset heterogeneity, algorithmic bias, lack of transparency, privacy concerns, medico-legal uncertainty, and automation bias still constrain broad imp
This narrative review, based on a structured search of PubMed/MEDLINE, Scopus, and Web of Science, maps the clinical applications of artificial intelligence across otolaryngology subspecialties, noting that deep learning shows potential in sinonasal disease detection, automated image segmentation, lymph node metastasis prediction, thyroid nodule classification, and prognostic modeling in head and neck cancer, while multimodal systems and generative large language models are emerging in medical education, image interpretation, differential diagnosis, and clinical decision support; however, limited external validation, retrospective designs, dataset heterogeneity, algorithmic bias, lack of transparency, privacy concerns, medico-legal uncertainty, and automation bias still constrain broad imp
This narrative review, based on a structured search of PubMed/MEDLINE, Scopus, and Web of Science, maps the clinical applications of artificial intelligence across otolaryngology subspecialties, noting that deep learning shows potential in sinonasal disease detection, automated image segmentation, lymph node metastasis prediction, thyroid nodule classification, and prognostic modeling in head and neck cancer, while multimodal systems and generative large language models are emerging in medical education, image interpretation, differential diagnosis, and clinical decision support; however, limited external validation, retrospective designs, dataset heterogeneity, algorithmic bias, lack of transparency, privacy concerns, medico-legal uncertainty, and automation bias still constrain broad imp
This narrative review, based on a structured search of PubMed/MEDLINE, Scopus, and Web of Science, maps the clinical applications of artificial intelligence across otolaryngology subspecialties, noting that deep learning shows potential in sinonasal disease detection, automated image segmentation, lymph node metastasis prediction, thyroid nodule classification, and prognostic modeling in head and neck cancer, while multimodal systems and generative large language models are emerging in medical education, image interpretation, differential diagnosis, and clinical decision support; however, limited external validation, retrospective designs, dataset heterogeneity, algorithmic bias, lack of transparency, privacy concerns, medico-legal uncertainty, and automation bias still constrain broad imp
bioRxiv Using paired transcriptomic and electrophysiological patch-sequencing data from 495 human neurons from neurosurgical tissue, this study compared how well traditional transcriptomic cell type classification, representations from the foundation model scGPT pretrained on large-scale scRNA-seq datasets, ion channel-coding genes, and highly variable genes predict electrophysiological features, finding that cluster-level cell type representations generally outperform highly variable gene selection, ion channel gene selection, and context-enriched scGPT embeddings, with the best results obtained by combining the outputs of separate cell type and scGPT-based models.
Using paired transcriptomic and electrophysiological patch-sequencing data from 495 human neurons from neurosurgical tissue, this study compared how well traditional transcriptomic cell type classification, representations from the foundation model scGPT pretrained on large-scale scRNA-seq datasets, ion channel-coding genes, and highly variable genes predict electrophysiological features, finding that cluster-level cell type representations generally outperform highly variable gene selection, ion channel gene selection, and context-enriched scGPT embeddings, with the best results obtained by combining the outputs of separate cell type and scGPT-based models.
Using paired transcriptomic and electrophysiological patch-sequencing data from 495 human neurons from neurosurgical tissue, this study compared how well traditional transcriptomic cell type classification, representations from the foundation model scGPT pretrained on large-scale scRNA-seq datasets, ion channel-coding genes, and highly variable genes predict electrophysiological features, finding that cluster-level cell type representations generally outperform highly variable gene selection, ion channel gene selection, and context-enriched scGPT embeddings, with the best results obtained by combining the outputs of separate cell type and scGPT-based models.
Using paired transcriptomic and electrophysiological patch-sequencing data from 495 human neurons from neurosurgical tissue, this study compared how well traditional transcriptomic cell type classification, representations from the foundation model scGPT pretrained on large-scale scRNA-seq datasets, ion channel-coding genes, and highly variable genes predict electrophysiological features, finding that cluster-level cell type representations generally outperform highly variable gene selection, ion channel gene selection, and context-enriched scGPT embeddings, with the best results obtained by combining the outputs of separate cell type and scGPT-based models.
medRxiv Using a zero-shot large language model pipeline (Gemini 3.5 Flash and Claude Haiku) to extract 46 symptoms from 2,728 discharge notes of 1,507 colorectal cancer patients in MIMIC-IV, this study built patient-level symptom co-occurrence networks with phi correlation (≥0.10) and Louvain community detection; both models converged on three clinically coherent symptom clusters — Systemic, CRC Disease-Specific, and Gastrointestinal — and Systemic cluster burden was associated with in-hospital mortality (OR=1.33) and 1-year mortality (OR=1.41) while CRC Disease-Specific cluster burden independently predicted 30-day readmission (OR=1.20), with both associations robust to adjustment for metastatic disease.
Using a zero-shot large language model pipeline (Gemini 3.5 Flash and Claude Haiku) to extract 46 symptoms from 2,728 discharge notes of 1,507 colorectal cancer patients in MIMIC-IV, this study built patient-level symptom co-occurrence networks with phi correlation (≥0.10) and Louvain community detection; both models converged on three clinically coherent symptom clusters — Systemic, CRC Disease-Specific, and Gastrointestinal — and Systemic cluster burden was associated with in-hospital mortality (OR=1.33) and 1-year mortality (OR=1.41) while CRC Disease-Specific cluster burden independently predicted 30-day readmission (OR=1.20), with both associations robust to adjustment for metastatic disease.
Using a zero-shot large language model pipeline (Gemini 3.5 Flash and Claude Haiku) to extract 46 symptoms from 2,728 discharge notes of 1,507 colorectal cancer patients in MIMIC-IV, this study built patient-level symptom co-occurrence networks with phi correlation (≥0.10) and Louvain community detection; both models converged on three clinically coherent symptom clusters — Systemic, CRC Disease-Specific, and Gastrointestinal — and Systemic cluster burden was associated with in-hospital mortality (OR=1.33) and 1-year mortality (OR=1.41) while CRC Disease-Specific cluster burden independently predicted 30-day readmission (OR=1.20), with both associations robust to adjustment for metastatic disease.
Using a zero-shot large language model pipeline (Gemini 3.5 Flash and Claude Haiku) to extract 46 symptoms from 2,728 discharge notes of 1,507 colorectal cancer patients in MIMIC-IV, this study built patient-level symptom co-occurrence networks with phi correlation (≥0.10) and Louvain community detection; both models converged on three clinically coherent symptom clusters — Systemic, CRC Disease-Specific, and Gastrointestinal — and Systemic cluster burden was associated with in-hospital mortality (OR=1.33) and 1-year mortality (OR=1.41) while CRC Disease-Specific cluster burden independently predicted 30-day readmission (OR=1.20), with both associations robust to adjustment for metastatic disease.
medRxiv Using choroid segmentation as a prototype task on 6,076 OCT B-scans from 80 subjects, this study systematically compared scalar no-reference image quality metrics (BRISQUE, NIQE, PIQE, SNR, PSNR), general-purpose ImageNet-pretrained representations, and retinal foundation models as task-specific quality gates, finding that scalar metrics correlate weakly with segmentation Dice (|r| < 0.20), that general-purpose pretrained representations reach linear-probe ROC-AUC up to about 0.77, that the OCT-specific foundation model RETFound reaches about 0.81, and that only the retinal foundation model embeddings form quality-aligned unsupervised K-Means clusters exceeding a patient-level permutation null.
Using choroid segmentation as a prototype task on 6,076 OCT B-scans from 80 subjects, this study systematically compared scalar no-reference image quality metrics (BRISQUE, NIQE, PIQE, SNR, PSNR), general-purpose ImageNet-pretrained representations, and retinal foundation models as task-specific quality gates, finding that scalar metrics correlate weakly with segmentation Dice (|r| < 0.20), that general-purpose pretrained representations reach linear-probe ROC-AUC up to about 0.77, that the OCT-specific foundation model RETFound reaches about 0.81, and that only the retinal foundation model embeddings form quality-aligned unsupervised K-Means clusters exceeding a patient-level permutation null.
Using choroid segmentation as a prototype task on 6,076 OCT B-scans from 80 subjects, this study systematically compared scalar no-reference image quality metrics (BRISQUE, NIQE, PIQE, SNR, PSNR), general-purpose ImageNet-pretrained representations, and retinal foundation models as task-specific quality gates, finding that scalar metrics correlate weakly with segmentation Dice (|r| < 0.20), that general-purpose pretrained representations reach linear-probe ROC-AUC up to about 0.77, that the OCT-specific foundation model RETFound reaches about 0.81, and that only the retinal foundation model embeddings form quality-aligned unsupervised K-Means clusters exceeding a patient-level permutation null.
Using choroid segmentation as a prototype task on 6,076 OCT B-scans from 80 subjects, this study systematically compared scalar no-reference image quality metrics (BRISQUE, NIQE, PIQE, SNR, PSNR), general-purpose ImageNet-pretrained representations, and retinal foundation models as task-specific quality gates, finding that scalar metrics correlate weakly with segmentation Dice (|r| < 0.20), that general-purpose pretrained representations reach linear-probe ROC-AUC up to about 0.77, that the OCT-specific foundation model RETFound reaches about 0.81, and that only the retinal foundation model embeddings form quality-aligned unsupervised K-Means clusters exceeding a patient-level permutation null.
RNA Biology The study proposes HyLnc, a framework that first pre-trains a custom BERT model on a large corpus of metazoan RNA sequences with masked language modelling, then fine-tunes it on curated lncRNA and protein-coding transcript datasets to extract 256-dimensional deep embeddings, while computing 348 handcrafted features including ORF characteristics, UTR properties, nucleotide composition and Fickett scores; after multi-stage feature selection, multiple machine learning classifiers were evaluated, with random forest performing best and achieving 91.30% accuracy, 91.23% F1-score and 82.60 MCC on an independent validation dataset, outperforming several existing lncRNA prediction tools.
The study proposes HyLnc, a framework that first pre-trains a custom BERT model on a large corpus of metazoan RNA sequences with masked language modelling, then fine-tunes it on curated lncRNA and protein-coding transcript datasets to extract 256-dimensional deep embeddings, while computing 348 handcrafted features including ORF characteristics, UTR properties, nucleotide composition and Fickett scores; after multi-stage feature selection, multiple machine learning classifiers were evaluated, with random forest performing best and achieving 91.30% accuracy, 91.23% F1-score and 82.60 MCC on an independent validation dataset, outperforming several existing lncRNA prediction tools.
The study proposes HyLnc, a framework that first pre-trains a custom BERT model on a large corpus of metazoan RNA sequences with masked language modelling, then fine-tunes it on curated lncRNA and protein-coding transcript datasets to extract 256-dimensional deep embeddings, while computing 348 handcrafted features including ORF characteristics, UTR properties, nucleotide composition and Fickett scores; after multi-stage feature selection, multiple machine learning classifiers were evaluated, with random forest performing best and achieving 91.30% accuracy, 91.23% F1-score and 82.60 MCC on an independent validation dataset, outperforming several existing lncRNA prediction tools.
The study proposes HyLnc, a framework that first pre-trains a custom BERT model on a large corpus of metazoan RNA sequences with masked language modelling, then fine-tunes it on curated lncRNA and protein-coding transcript datasets to extract 256-dimensional deep embeddings, while computing 348 handcrafted features including ORF characteristics, UTR properties, nucleotide composition and Fickett scores; after multi-stage feature selection, multiple machine learning classifiers were evaluated, with random forest performing best and achieving 91.30% accuracy, 91.23% F1-score and 82.60 MCC on an independent validation dataset, outperforming several existing lncRNA prediction tools.
Aesthetic Surgery Journal This systematic review searched studies from 2000 to 2025 evaluating aesthetic outcomes after implant-based, autologous, or hybrid breast reconstruction, included 51 studies with 7711 participants from 16 countries, classified assessments as subjective or objective, found subjective tools led by BREAST-Q (35/51, 69%) and objective methods in 19 studies including 3-dimensional surface imaging (8/51, 16%), BCCT.core (7/51, 14%), eye tracking (3/51, 6%), and artificial intelligence-based analyses (3/51, 6%), and noted that no single modality comprehensively addresses all aesthetic domains, with a progressive shift toward multimodal evaluation.
This systematic review searched studies from 2000 to 2025 evaluating aesthetic outcomes after implant-based, autologous, or hybrid breast reconstruction, included 51 studies with 7711 participants from 16 countries, classified assessments as subjective or objective, found subjective tools led by BREAST-Q (35/51, 69%) and objective methods in 19 studies including 3-dimensional surface imaging (8/51, 16%), BCCT.core (7/51, 14%), eye tracking (3/51, 6%), and artificial intelligence-based analyses (3/51, 6%), and noted that no single modality comprehensively addresses all aesthetic domains, with a progressive shift toward multimodal evaluation.
This systematic review searched studies from 2000 to 2025 evaluating aesthetic outcomes after implant-based, autologous, or hybrid breast reconstruction, included 51 studies with 7711 participants from 16 countries, classified assessments as subjective or objective, found subjective tools led by BREAST-Q (35/51, 69%) and objective methods in 19 studies including 3-dimensional surface imaging (8/51, 16%), BCCT.core (7/51, 14%), eye tracking (3/51, 6%), and artificial intelligence-based analyses (3/51, 6%), and noted that no single modality comprehensively addresses all aesthetic domains, with a progressive shift toward multimodal evaluation.
This systematic review searched studies from 2000 to 2025 evaluating aesthetic outcomes after implant-based, autologous, or hybrid breast reconstruction, included 51 studies with 7711 participants from 16 countries, classified assessments as subjective or objective, found subjective tools led by BREAST-Q (35/51, 69%) and objective methods in 19 studies including 3-dimensional surface imaging (8/51, 16%), BCCT.core (7/51, 14%), eye tracking (3/51, 6%), and artificial intelligence-based analyses (3/51, 6%), and noted that no single modality comprehensively addresses all aesthetic domains, with a progressive shift toward multimodal evaluation.
bioRxiv This work presents nanorepertoire, an end-to-end Nextflow DSL2 pipeline dedicated to camelid VHH nanobody repertoires that takes paired-end AIRR-seq FASTQ files through quality control, adapter trimming, read merging, in-silico translation, CD-HIT clonotyping and deep-learning CDR3 annotation with nanoCDR-X, and returns an interactive HTML report covering clonal architecture, CDR3 length and amino-acid composition, intra-clonal homogeneity, repertoire diversity and the computational carbon footprint of the run; applied to two publicly available SARS-CoV-2 RBD-selected llama libraries (4.
This work presents nanorepertoire, an end-to-end Nextflow DSL2 pipeline dedicated to camelid VHH nanobody repertoires that takes paired-end AIRR-seq FASTQ files through quality control, adapter trimming, read merging, in-silico translation, CD-HIT clonotyping and deep-learning CDR3 annotation with nanoCDR-X, and returns an interactive HTML report covering clonal architecture, CDR3 length and amino-acid composition, intra-clonal homogeneity, repertoire diversity and the computational carbon footprint of the run; applied to two publicly available SARS-CoV-2 RBD-selected llama libraries (4.
This work presents nanorepertoire, an end-to-end Nextflow DSL2 pipeline dedicated to camelid VHH nanobody repertoires that takes paired-end AIRR-seq FASTQ files through quality control, adapter trimming, read merging, in-silico translation, CD-HIT clonotyping and deep-learning CDR3 annotation with nanoCDR-X, and returns an interactive HTML report covering clonal architecture, CDR3 length and amino-acid composition, intra-clonal homogeneity, repertoire diversity and the computational carbon footprint of the run; applied to two publicly available SARS-CoV-2 RBD-selected llama libraries (4.
This work presents nanorepertoire, an end-to-end Nextflow DSL2 pipeline dedicated to camelid VHH nanobody repertoires that takes paired-end AIRR-seq FASTQ files through quality control, adapter trimming, read merging, in-silico translation, CD-HIT clonotyping and deep-learning CDR3 annotation with nanoCDR-X, and returns an interactive HTML report covering clonal architecture, CDR3 length and amino-acid composition, intra-clonal homogeneity, repertoire diversity and the computational carbon footprint of the run; applied to two publicly available SARS-CoV-2 RBD-selected llama libraries (4.
Heart & Lung This study retrospectively analyzed 615 ECGs from 270 emergency medical service patient encounters with concern for acute myocardial infarction, using cardiology activation of the STEMI pathway as the reference standard, and compared three multimodal large language models with an ECG machine algorithm, finding that Gemini had the highest sensitivity (95.3%) but extremely poor specificity (9.4%), ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%), while the ECG machine algorithm was more balanced with sensitivity of 67.7% and specificity of 64.2%, suggesting that general-purpose large language models are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive emergency workflows.
This study retrospectively analyzed 615 ECGs from 270 emergency medical service patient encounters with concern for acute myocardial infarction, using cardiology activation of the STEMI pathway as the reference standard, and compared three multimodal large language models with an ECG machine algorithm, finding that Gemini had the highest sensitivity (95.3%) but extremely poor specificity (9.4%), ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%), while the ECG machine algorithm was more balanced with sensitivity of 67.7% and specificity of 64.2%, suggesting that general-purpose large language models are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive emergency workflows.
This study retrospectively analyzed 615 ECGs from 270 emergency medical service patient encounters with concern for acute myocardial infarction, using cardiology activation of the STEMI pathway as the reference standard, and compared three multimodal large language models with an ECG machine algorithm, finding that Gemini had the highest sensitivity (95.3%) but extremely poor specificity (9.4%), ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%), while the ECG machine algorithm was more balanced with sensitivity of 67.7% and specificity of 64.2%, suggesting that general-purpose large language models are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive emergency workflows.
This study retrospectively analyzed 615 ECGs from 270 emergency medical service patient encounters with concern for acute myocardial infarction, using cardiology activation of the STEMI pathway as the reference standard, and compared three multimodal large language models with an ECG machine algorithm, finding that Gemini had the highest sensitivity (95.3%) but extremely poor specificity (9.4%), ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%), while the ECG machine algorithm was more balanced with sensitivity of 67.7% and specificity of 64.2%, suggesting that general-purpose large language models are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive emergency workflows.
bioRxiv This work develops a method to infer a Corpus-Wide Causal Score (CWCS) for a gene-disease pair by integrating network-based causal signals in a gene regulatory network (CWCS-Net) with corpus-wide literature evidence from PubMed abstracts quantified by a newly developed Truth Discovery algorithm (CWCS-TD), achieving a causal class F1 score of 0.600 across ten diseases using OMIM as an external expert-curated reference, outperforming GPT-4o (0.505) and MMed-Llama 3 (0.522).
This work develops a method to infer a Corpus-Wide Causal Score (CWCS) for a gene-disease pair by integrating network-based causal signals in a gene regulatory network (CWCS-Net) with corpus-wide literature evidence from PubMed abstracts quantified by a newly developed Truth Discovery algorithm (CWCS-TD), achieving a causal class F1 score of 0.600 across ten diseases using OMIM as an external expert-curated reference, outperforming GPT-4o (0.505) and MMed-Llama 3 (0.522).
This work develops a method to infer a Corpus-Wide Causal Score (CWCS) for a gene-disease pair by integrating network-based causal signals in a gene regulatory network (CWCS-Net) with corpus-wide literature evidence from PubMed abstracts quantified by a newly developed Truth Discovery algorithm (CWCS-TD), achieving a causal class F1 score of 0.600 across ten diseases using OMIM as an external expert-curated reference, outperforming GPT-4o (0.505) and MMed-Llama 3 (0.522).
This work develops a method to infer a Corpus-Wide Causal Score (CWCS) for a gene-disease pair by integrating network-based causal signals in a gene regulatory network (CWCS-Net) with corpus-wide literature evidence from PubMed abstracts quantified by a newly developed Truth Discovery algorithm (CWCS-TD), achieving a causal class F1 score of 0.600 across ten diseases using OMIM as an external expert-curated reference, outperforming GPT-4o (0.505) and MMed-Llama 3 (0.522).
medRxiv Using harmonised clinical data from two Phase 3 trials (2,918 participants), this study trained monthly tabular models from baseline to therapy end for time-resolved prediction of end-of-therapy (EOT) outcomes and post-treatment relapse, finding that EOT outcome prediction improved after month 3 (ROC-AUC up to 0.84, driven by sputum-smear and solid culture), whereas relapse prediction among those with favourable EOT outcomes and completed follow-up improved only modestly through month 3 (ROC-AUC 0.58-0.63) before declining, with age, sex, clinical symptoms and bacterial burden contributing most; large language model-derived embedding models matched tabular relapse models throughout therapy and outperformed tabular EOT models at months 3 and 4 (ΔROC-AUC 0.12 and 0.
Using harmonised clinical data from two Phase 3 trials (2,918 participants), this study trained monthly tabular models from baseline to therapy end for time-resolved prediction of end-of-therapy (EOT) outcomes and post-treatment relapse, finding that EOT outcome prediction improved after month 3 (ROC-AUC up to 0.84, driven by sputum-smear and solid culture), whereas relapse prediction among those with favourable EOT outcomes and completed follow-up improved only modestly through month 3 (ROC-AUC 0.58-0.63) before declining, with age, sex, clinical symptoms and bacterial burden contributing most; large language model-derived embedding models matched tabular relapse models throughout therapy and outperformed tabular EOT models at months 3 and 4 (ΔROC-AUC 0.12 and 0.
Using harmonised clinical data from two Phase 3 trials (2,918 participants), this study trained monthly tabular models from baseline to therapy end for time-resolved prediction of end-of-therapy (EOT) outcomes and post-treatment relapse, finding that EOT outcome prediction improved after month 3 (ROC-AUC up to 0.84, driven by sputum-smear and solid culture), whereas relapse prediction among those with favourable EOT outcomes and completed follow-up improved only modestly through month 3 (ROC-AUC 0.58-0.63) before declining, with age, sex, clinical symptoms and bacterial burden contributing most; large language model-derived embedding models matched tabular relapse models throughout therapy and outperformed tabular EOT models at months 3 and 4 (ΔROC-AUC 0.12 and 0.
Using harmonised clinical data from two Phase 3 trials (2,918 participants), this study trained monthly tabular models from baseline to therapy end for time-resolved prediction of end-of-therapy (EOT) outcomes and post-treatment relapse, finding that EOT outcome prediction improved after month 3 (ROC-AUC up to 0.84, driven by sputum-smear and solid culture), whereas relapse prediction among those with favourable EOT outcomes and completed follow-up improved only modestly through month 3 (ROC-AUC 0.58-0.63) before declining, with age, sex, clinical symptoms and bacterial burden contributing most; large language model-derived embedding models matched tabular relapse models throughout therapy and outperformed tabular EOT models at months 3 and 4 (ΔROC-AUC 0.12 and 0.
Pharmacology & therapeutics This review proposes reframing multidrug resistance (MDR) from conventional single-gene mechanism models toward systemic adaptive resistance as a complex and evolving biological phenomenon, and systematically reviews classical resistance mechanisms (drug uptake/efflux, compensatory pathway activation, apoptosis evasion), non-genetic resistance mechanisms (drug-tolerant persisters, DTPs; epithelial-mesenchymal transition, EMT; cancer stem cell, CSC, plasticity and related factors), how interactions among these mechanisms contribute to tumor persistence and MDR evolution, and translational barriers together with emerging pharmacological strategies including longitudinal molecular monitoring and artificial intelligence (AI)-assisted drug resistance prediction.
This review proposes reframing multidrug resistance (MDR) from conventional single-gene mechanism models toward systemic adaptive resistance as a complex and evolving biological phenomenon, and systematically reviews classical resistance mechanisms (drug uptake/efflux, compensatory pathway activation, apoptosis evasion), non-genetic resistance mechanisms (drug-tolerant persisters, DTPs; epithelial-mesenchymal transition, EMT; cancer stem cell, CSC, plasticity and related factors), how interactions among these mechanisms contribute to tumor persistence and MDR evolution, and translational barriers together with emerging pharmacological strategies including longitudinal molecular monitoring and artificial intelligence (AI)-assisted drug resistance prediction.
This review proposes reframing multidrug resistance (MDR) from conventional single-gene mechanism models toward systemic adaptive resistance as a complex and evolving biological phenomenon, and systematically reviews classical resistance mechanisms (drug uptake/efflux, compensatory pathway activation, apoptosis evasion), non-genetic resistance mechanisms (drug-tolerant persisters, DTPs; epithelial-mesenchymal transition, EMT; cancer stem cell, CSC, plasticity and related factors), how interactions among these mechanisms contribute to tumor persistence and MDR evolution, and translational barriers together with emerging pharmacological strategies including longitudinal molecular monitoring and artificial intelligence (AI)-assisted drug resistance prediction.
This review proposes reframing multidrug resistance (MDR) from conventional single-gene mechanism models toward systemic adaptive resistance as a complex and evolving biological phenomenon, and systematically reviews classical resistance mechanisms (drug uptake/efflux, compensatory pathway activation, apoptosis evasion), non-genetic resistance mechanisms (drug-tolerant persisters, DTPs; epithelial-mesenchymal transition, EMT; cancer stem cell, CSC, plasticity and related factors), how interactions among these mechanisms contribute to tumor persistence and MDR evolution, and translational barriers together with emerging pharmacological strategies including longitudinal molecular monitoring and artificial intelligence (AI)-assisted drug resistance prediction.
medRxiv This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.
This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.
This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.
This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.
arXiv The work proposes dynamic context adaptation: a validation-generation loop in which a validation agent extracts structured diagnostic feedback from execution traces, a generation agent proposes multiple candidates per iteration, a knowledge graph supplies semantic constraints, and simulated annealing performs non-greedy selection, targeting the LLM code-generation limitation the authors call static binding; across eight problems the method outperforms zero-shot, Reflexion, and OpenEvolve on seven of eight problems at both 300 and 600 evaluations (p < 0.01), achieves the best score at 1000 evaluations on the primary motivating problem of cross-coupled optimization (0.694 vs. 0.681, p = 0.019, d = 0.52), and ablations identify structured execution feedback as the primary driver.
The work proposes dynamic context adaptation: a validation-generation loop in which a validation agent extracts structured diagnostic feedback from execution traces, a generation agent proposes multiple candidates per iteration, a knowledge graph supplies semantic constraints, and simulated annealing performs non-greedy selection, targeting the LLM code-generation limitation the authors call static binding; across eight problems the method outperforms zero-shot, Reflexion, and OpenEvolve on seven of eight problems at both 300 and 600 evaluations (p < 0.01), achieves the best score at 1000 evaluations on the primary motivating problem of cross-coupled optimization (0.694 vs. 0.681, p = 0.019, d = 0.52), and ablations identify structured execution feedback as the primary driver.
The work proposes dynamic context adaptation: a validation-generation loop in which a validation agent extracts structured diagnostic feedback from execution traces, a generation agent proposes multiple candidates per iteration, a knowledge graph supplies semantic constraints, and simulated annealing performs non-greedy selection, targeting the LLM code-generation limitation the authors call static binding; across eight problems the method outperforms zero-shot, Reflexion, and OpenEvolve on seven of eight problems at both 300 and 600 evaluations (p < 0.01), achieves the best score at 1000 evaluations on the primary motivating problem of cross-coupled optimization (0.694 vs. 0.681, p = 0.019, d = 0.52), and ablations identify structured execution feedback as the primary driver.
The work proposes dynamic context adaptation: a validation-generation loop in which a validation agent extracts structured diagnostic feedback from execution traces, a generation agent proposes multiple candidates per iteration, a knowledge graph supplies semantic constraints, and simulated annealing performs non-greedy selection, targeting the LLM code-generation limitation the authors call static binding; across eight problems the method outperforms zero-shot, Reflexion, and OpenEvolve on seven of eight problems at both 300 and 600 evaluations (p < 0.01), achieves the best score at 1000 evaluations on the primary motivating problem of cross-coupled optimization (0.694 vs. 0.681, p = 0.019, d = 0.52), and ablations identify structured execution feedback as the primary driver.
arXiv This work proposes GenFT, a W0-conditioned parameter-efficient fine-tuning method in which a deterministic generator produces task-specific updates ΔW by applying row and column transformations to the pretrained weights W0 together with a shared-specific decomposition, reporting competitive or better average performance on GLUE (RoBERTaBase, 85.87% average with 0.24M parameters), VTAB-1K (ViT-B/16, 74.50% average with 0.27M parameters) and FGVC (90.38% average with 0.29M parameters), plus a perplexity pilot study on LLaMA-7B with an Alpaca subset.
This work proposes GenFT, a W0-conditioned parameter-efficient fine-tuning method in which a deterministic generator produces task-specific updates ΔW by applying row and column transformations to the pretrained weights W0 together with a shared-specific decomposition, reporting competitive or better average performance on GLUE (RoBERTaBase, 85.87% average with 0.24M parameters), VTAB-1K (ViT-B/16, 74.50% average with 0.27M parameters) and FGVC (90.38% average with 0.29M parameters), plus a perplexity pilot study on LLaMA-7B with an Alpaca subset.
This work proposes GenFT, a W0-conditioned parameter-efficient fine-tuning method in which a deterministic generator produces task-specific updates ΔW by applying row and column transformations to the pretrained weights W0 together with a shared-specific decomposition, reporting competitive or better average performance on GLUE (RoBERTaBase, 85.87% average with 0.24M parameters), VTAB-1K (ViT-B/16, 74.50% average with 0.27M parameters) and FGVC (90.38% average with 0.29M parameters), plus a perplexity pilot study on LLaMA-7B with an Alpaca subset.
This work proposes GenFT, a W0-conditioned parameter-efficient fine-tuning method in which a deterministic generator produces task-specific updates ΔW by applying row and column transformations to the pretrained weights W0 together with a shared-specific decomposition, reporting competitive or better average performance on GLUE (RoBERTaBase, 85.87% average with 0.24M parameters), VTAB-1K (ViT-B/16, 74.50% average with 0.27M parameters) and FGVC (90.38% average with 0.29M parameters), plus a perplexity pilot study on LLaMA-7B with an Alpaca subset.
Journal of the European Meteorological Society. This article by Peter Dueben, Peter Bauer, Oliver Fuhrer, Nikolay Koldunov and Jørn Kristiansen argues that, following the success of machine learning in producing weather predictions with competitive skill compared to complex traditional systems, attention should shift from forecast output to the working practices that make prediction systems possible, and that machine learning and recent digital technologies will reshape the forecasting value chain — how models are coded and developed, how observations and Earth-system data are exploited, how data and computing are managed, how systems are verified, and how information is created, evaluated and turned into services; it discusses six non-exhaustive areas in which agentic software engineering, open and compressed data, shared verification
This article by Peter Dueben, Peter Bauer, Oliver Fuhrer, Nikolay Koldunov and Jørn Kristiansen argues that, following the success of machine learning in producing weather predictions with competitive skill compared to complex traditional systems, attention should shift from forecast output to the working practices that make prediction systems possible, and that machine learning and recent digital technologies will reshape the forecasting value chain — how models are coded and developed, how observations and Earth-system data are exploited, how data and computing are managed, how systems are verified, and how information is created, evaluated and turned into services; it discusses six non-exhaustive areas in which agentic software engineering, open and compressed data, shared verification
This article by Peter Dueben, Peter Bauer, Oliver Fuhrer, Nikolay Koldunov and Jørn Kristiansen argues that, following the success of machine learning in producing weather predictions with competitive skill compared to complex traditional systems, attention should shift from forecast output to the working practices that make prediction systems possible, and that machine learning and recent digital technologies will reshape the forecasting value chain — how models are coded and developed, how observations and Earth-system data are exploited, how data and computing are managed, how systems are verified, and how information is created, evaluated and turned into services; it discusses six non-exhaustive areas in which agentic software engineering, open and compressed data, shared verification
This article by Peter Dueben, Peter Bauer, Oliver Fuhrer, Nikolay Koldunov and Jørn Kristiansen argues that, following the success of machine learning in producing weather predictions with competitive skill compared to complex traditional systems, attention should shift from forecast output to the working practices that make prediction systems possible, and that machine learning and recent digital technologies will reshape the forecasting value chain — how models are coded and developed, how observations and Earth-system data are exploited, how data and computing are managed, how systems are verified, and how information is created, evaluated and turned into services; it discusses six non-exhaustive areas in which agentic software engineering, open and compressed data, shared verification
Journal of Medical Internet Research Using electronic health record data from 9,496 patients with trauma treated in the Capital Region of Denmark between 2017 and 2024, this study developed a hybrid neural network combining tabular and sequential data to predict 30-day all-cause mortality at any time point from prehospital care to discharge, achieving AUROC 0.962 and AUPRC 0.655 on a holdout set of 1,829 patients and AUROC 0.905 at 1 hour from first patient contact in active-cohort evaluation, with better discrimination than the Revised Trauma Score (mean ΔAUROC +0.297) and the Trauma and Injury Severity Score (mean ΔAUROC +0.169).
Using electronic health record data from 9,496 patients with trauma treated in the Capital Region of Denmark between 2017 and 2024, this study developed a hybrid neural network combining tabular and sequential data to predict 30-day all-cause mortality at any time point from prehospital care to discharge, achieving AUROC 0.962 and AUPRC 0.655 on a holdout set of 1,829 patients and AUROC 0.905 at 1 hour from first patient contact in active-cohort evaluation, with better discrimination than the Revised Trauma Score (mean ΔAUROC +0.297) and the Trauma and Injury Severity Score (mean ΔAUROC +0.169).
Using electronic health record data from 9,496 patients with trauma treated in the Capital Region of Denmark between 2017 and 2024, this study developed a hybrid neural network combining tabular and sequential data to predict 30-day all-cause mortality at any time point from prehospital care to discharge, achieving AUROC 0.962 and AUPRC 0.655 on a holdout set of 1,829 patients and AUROC 0.905 at 1 hour from first patient contact in active-cohort evaluation, with better discrimination than the Revised Trauma Score (mean ΔAUROC +0.297) and the Trauma and Injury Severity Score (mean ΔAUROC +0.169).
Using electronic health record data from 9,496 patients with trauma treated in the Capital Region of Denmark between 2017 and 2024, this study developed a hybrid neural network combining tabular and sequential data to predict 30-day all-cause mortality at any time point from prehospital care to discharge, achieving AUROC 0.962 and AUPRC 0.655 on a holdout set of 1,829 patients and AUROC 0.905 at 1 hour from first patient contact in active-cohort evaluation, with better discrimination than the Revised Trauma Score (mean ΔAUROC +0.297) and the Trauma and Injury Severity Score (mean ΔAUROC +0.169).
Transfusion clinique et biologique : journal de la Societe francaise de transfusion sanguine This review reflects on how infectious diseases shaped transfusion services from 2016 to 2026, covering innovations in donor selection, testing and pathogen reduction, notable pathogens including Zika virus, Plasmodium, Babesia and SARS-CoV-2, and favorable developments such as individualized risk assessment, relaxation of donor deferral policies, and AI and machine learning tools for surveillance and horizon scanning, while noting that many of these gains do not extend to low and low-middle income countries.
This review reflects on how infectious diseases shaped transfusion services from 2016 to 2026, covering innovations in donor selection, testing and pathogen reduction, notable pathogens including Zika virus, Plasmodium, Babesia and SARS-CoV-2, and favorable developments such as individualized risk assessment, relaxation of donor deferral policies, and AI and machine learning tools for surveillance and horizon scanning, while noting that many of these gains do not extend to low and low-middle income countries.
This review reflects on how infectious diseases shaped transfusion services from 2016 to 2026, covering innovations in donor selection, testing and pathogen reduction, notable pathogens including Zika virus, Plasmodium, Babesia and SARS-CoV-2, and favorable developments such as individualized risk assessment, relaxation of donor deferral policies, and AI and machine learning tools for surveillance and horizon scanning, while noting that many of these gains do not extend to low and low-middle income countries.
This review reflects on how infectious diseases shaped transfusion services from 2016 to 2026, covering innovations in donor selection, testing and pathogen reduction, notable pathogens including Zika virus, Plasmodium, Babesia and SARS-CoV-2, and favorable developments such as individualized risk assessment, relaxation of donor deferral policies, and AI and machine learning tools for surveillance and horizon scanning, while noting that many of these gains do not extend to low and low-middle income countries.
International Journal of Cardiology OPTIMUST was a prospective, single-center implementation study in which physicians without formal echocardiography certification completed a structured two-month curriculum and performed focused examinations using Caption AI-enabled handheld ultrasound on cardiology and non-cardiology wards; of 287 attempted examinations, 206 (71.8%) were analyzable, operator assessment correlated with expert review of the same handheld image sets for LVEF (r = 0.84), filling-pressure category agreement gave a quadratic weighted κ of 0.659, physicians reported that handheld findings changed or confirmed management in 93.7% of examinations, and 32.8% underwent comprehensive echocardiography within one month.
OPTIMUST was a prospective, single-center implementation study in which physicians without formal echocardiography certification completed a structured two-month curriculum and performed focused examinations using Caption AI-enabled handheld ultrasound on cardiology and non-cardiology wards; of 287 attempted examinations, 206 (71.8%) were analyzable, operator assessment correlated with expert review of the same handheld image sets for LVEF (r = 0.84), filling-pressure category agreement gave a quadratic weighted κ of 0.659, physicians reported that handheld findings changed or confirmed management in 93.7% of examinations, and 32.8% underwent comprehensive echocardiography within one month.
OPTIMUST was a prospective, single-center implementation study in which physicians without formal echocardiography certification completed a structured two-month curriculum and performed focused examinations using Caption AI-enabled handheld ultrasound on cardiology and non-cardiology wards; of 287 attempted examinations, 206 (71.8%) were analyzable, operator assessment correlated with expert review of the same handheld image sets for LVEF (r = 0.84), filling-pressure category agreement gave a quadratic weighted κ of 0.659, physicians reported that handheld findings changed or confirmed management in 93.7% of examinations, and 32.8% underwent comprehensive echocardiography within one month.
OPTIMUST was a prospective, single-center implementation study in which physicians without formal echocardiography certification completed a structured two-month curriculum and performed focused examinations using Caption AI-enabled handheld ultrasound on cardiology and non-cardiology wards; of 287 attempted examinations, 206 (71.8%) were analyzable, operator assessment correlated with expert review of the same handheld image sets for LVEF (r = 0.84), filling-pressure category agreement gave a quadratic weighted κ of 0.659, physicians reported that handheld findings changed or confirmed management in 93.7% of examinations, and 32.8% underwent comprehensive echocardiography within one month.
medRxiv Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.
Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.
Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.
Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.
Clinical transplantation and research This review synthesizes current evidence on the use of artificial intelligence and machine learning in preclinical nonhuman primate xenotransplantation, noting that computer vision can support continuous noninvasive behavioral phenotyping and pain assessment, anomaly detection algorithms can extract early warning signals from biosignal streams, multiomics integration can identify xenograft-specific biomarker signatures, digital twin frameworks may enable hypothesis-generating simulations of prospective human recipient responses, and explainable AI can make model outputs more transparent and regulatory defensible, and it proposes a research agenda through which AI-augmented nonhuman primate experimentation can become a bridge from preclinical data complexity to clinical precision in xenotra
This review synthesizes current evidence on the use of artificial intelligence and machine learning in preclinical nonhuman primate xenotransplantation, noting that computer vision can support continuous noninvasive behavioral phenotyping and pain assessment, anomaly detection algorithms can extract early warning signals from biosignal streams, multiomics integration can identify xenograft-specific biomarker signatures, digital twin frameworks may enable hypothesis-generating simulations of prospective human recipient responses, and explainable AI can make model outputs more transparent and regulatory defensible, and it proposes a research agenda through which AI-augmented nonhuman primate experimentation can become a bridge from preclinical data complexity to clinical precision in xenotra
This review synthesizes current evidence on the use of artificial intelligence and machine learning in preclinical nonhuman primate xenotransplantation, noting that computer vision can support continuous noninvasive behavioral phenotyping and pain assessment, anomaly detection algorithms can extract early warning signals from biosignal streams, multiomics integration can identify xenograft-specific biomarker signatures, digital twin frameworks may enable hypothesis-generating simulations of prospective human recipient responses, and explainable AI can make model outputs more transparent and regulatory defensible, and it proposes a research agenda through which AI-augmented nonhuman primate experimentation can become a bridge from preclinical data complexity to clinical precision in xenotra
This review synthesizes current evidence on the use of artificial intelligence and machine learning in preclinical nonhuman primate xenotransplantation, noting that computer vision can support continuous noninvasive behavioral phenotyping and pain assessment, anomaly detection algorithms can extract early warning signals from biosignal streams, multiomics integration can identify xenograft-specific biomarker signatures, digital twin frameworks may enable hypothesis-generating simulations of prospective human recipient responses, and explainable AI can make model outputs more transparent and regulatory defensible, and it proposes a research agenda through which AI-augmented nonhuman primate experimentation can become a bridge from preclinical data complexity to clinical precision in xenotra
发表出处待核验 In the context of family caregivers of people living with Alzheimer's Disease and Related Dementias (ADRD), this work compares caregiver support exchanges from online communities with peer-like responses prompted from three LLMs (LLaMA, GPT-4o-mini, and MedGemma), using psycholinguistic and qualitative analysis to show that peer responses used significantly more first-person and past-focused language than peer-like AI responses, identifies seven types of personal narratives in human peer support, and finds that AI often captures their emotional work while potentially fabricating experiential grounding, thereby naming a narrative authenticity gap and a synthetic lived experience paradox.
In the context of family caregivers of people living with Alzheimer's Disease and Related Dementias (ADRD), this work compares caregiver support exchanges from online communities with peer-like responses prompted from three LLMs (LLaMA, GPT-4o-mini, and MedGemma), using psycholinguistic and qualitative analysis to show that peer responses used significantly more first-person and past-focused language than peer-like AI responses, identifies seven types of personal narratives in human peer support, and finds that AI often captures their emotional work while potentially fabricating experiential grounding, thereby naming a narrative authenticity gap and a synthetic lived experience paradox.
In the context of family caregivers of people living with Alzheimer's Disease and Related Dementias (ADRD), this work compares caregiver support exchanges from online communities with peer-like responses prompted from three LLMs (LLaMA, GPT-4o-mini, and MedGemma), using psycholinguistic and qualitative analysis to show that peer responses used significantly more first-person and past-focused language than peer-like AI responses, identifies seven types of personal narratives in human peer support, and finds that AI often captures their emotional work while potentially fabricating experiential grounding, thereby naming a narrative authenticity gap and a synthetic lived experience paradox.
In the context of family caregivers of people living with Alzheimer's Disease and Related Dementias (ADRD), this work compares caregiver support exchanges from online communities with peer-like responses prompted from three LLMs (LLaMA, GPT-4o-mini, and MedGemma), using psycholinguistic and qualitative analysis to show that peer responses used significantly more first-person and past-focused language than peer-like AI responses, identifies seven types of personal narratives in human peer support, and finds that AI often captures their emotional work while potentially fabricating experiential grounding, thereby naming a narrative authenticity gap and a synthetic lived experience paradox.
arXiv This work proposes a new self-supervised pretraining task for 3D graph neural networks: pretrained on 542k SwissProt protein structures from the AlphaFold Database, the model predicts the Euclidean distance between the geometric centroid of each protein subgraph (2-hop ego networks centered on 10% of amino acids) and the global geometric centroid of the whole protein, discretized into 10 equal bins and trained with cross-entropy; across ProNet, SchNet, and GCN backbones and ca_base, ca_angles, and ca_bb featurizations, it improves Fold, Superfamily, Family, and React classification by up to about 6% over no-pretraining and edge-distance pretraining baselines, without multiple views, augmentations, or masking strategies.
This work proposes a new self-supervised pretraining task for 3D graph neural networks: pretrained on 542k SwissProt protein structures from the AlphaFold Database, the model predicts the Euclidean distance between the geometric centroid of each protein subgraph (2-hop ego networks centered on 10% of amino acids) and the global geometric centroid of the whole protein, discretized into 10 equal bins and trained with cross-entropy; across ProNet, SchNet, and GCN backbones and ca_base, ca_angles, and ca_bb featurizations, it improves Fold, Superfamily, Family, and React classification by up to about 6% over no-pretraining and edge-distance pretraining baselines, without multiple views, augmentations, or masking strategies.
This work proposes a new self-supervised pretraining task for 3D graph neural networks: pretrained on 542k SwissProt protein structures from the AlphaFold Database, the model predicts the Euclidean distance between the geometric centroid of each protein subgraph (2-hop ego networks centered on 10% of amino acids) and the global geometric centroid of the whole protein, discretized into 10 equal bins and trained with cross-entropy; across ProNet, SchNet, and GCN backbones and ca_base, ca_angles, and ca_bb featurizations, it improves Fold, Superfamily, Family, and React classification by up to about 6% over no-pretraining and edge-distance pretraining baselines, without multiple views, augmentations, or masking strategies.
This work proposes a new self-supervised pretraining task for 3D graph neural networks: pretrained on 542k SwissProt protein structures from the AlphaFold Database, the model predicts the Euclidean distance between the geometric centroid of each protein subgraph (2-hop ego networks centered on 10% of amino acids) and the global geometric centroid of the whole protein, discretized into 10 equal bins and trained with cross-entropy; across ProNet, SchNet, and GCN backbones and ca_base, ca_angles, and ca_bb featurizations, it improves Fold, Superfamily, Family, and React classification by up to about 6% over no-pretraining and edge-distance pretraining baselines, without multiple views, augmentations, or masking strategies.
The Journal of Clinical Endocrinology & Metabolism This study developed and deployed a transformer-based natural language processing pipeline to identify incidental thyroid findings in radiology reports from 115,683 adults without prior thyroid disease across Mayo Clinic sites from July 1, 2017, to September 30, 2023, finding that 7.8% had such findings (92.9% nodular) and that these findings were associated with higher odds of downstream thyroid nodule diagnosis, biopsy, thyroidectomy, and thyroid cancer diagnosis, with most cancers being papillary.
This study developed and deployed a transformer-based natural language processing pipeline to identify incidental thyroid findings in radiology reports from 115,683 adults without prior thyroid disease across Mayo Clinic sites from July 1, 2017, to September 30, 2023, finding that 7.8% had such findings (92.9% nodular) and that these findings were associated with higher odds of downstream thyroid nodule diagnosis, biopsy, thyroidectomy, and thyroid cancer diagnosis, with most cancers being papillary.
This study developed and deployed a transformer-based natural language processing pipeline to identify incidental thyroid findings in radiology reports from 115,683 adults without prior thyroid disease across Mayo Clinic sites from July 1, 2017, to September 30, 2023, finding that 7.8% had such findings (92.9% nodular) and that these findings were associated with higher odds of downstream thyroid nodule diagnosis, biopsy, thyroidectomy, and thyroid cancer diagnosis, with most cancers being papillary.
This study developed and deployed a transformer-based natural language processing pipeline to identify incidental thyroid findings in radiology reports from 115,683 adults without prior thyroid disease across Mayo Clinic sites from July 1, 2017, to September 30, 2023, finding that 7.8% had such findings (92.9% nodular) and that these findings were associated with higher odds of downstream thyroid nodule diagnosis, biopsy, thyroidectomy, and thyroid cancer diagnosis, with most cancers being papillary.
Trends in biotechnology This review focuses on liquid biopsy for glioblastoma, noting that tissue biopsy is invasive and fails to capture tumor heterogeneity or temporal dynamics, while liquid biopsy can provide noninvasive real-time monitoring via circulating biomarkers such as cell-free DNA, circulating tumor DNA, circulating tumor cells, and extracellular vesicles in plasma, cerebrospinal fluid, urine, and saliva; it advances beyond prior biomarker catalogs by delivering a quantitative technology scorecard comparing cfDNA-, CTC-, and EV-based platforms in terms of sensitivity, clinical actionability, cost, and scalability, and introduces a multimodal decision matrix and an AI-driven fusion pipeline integrating fragmentomics, EV proteomics, and CTC transcriptomics to enhance minimal residual disease detection a
This review focuses on liquid biopsy for glioblastoma, noting that tissue biopsy is invasive and fails to capture tumor heterogeneity or temporal dynamics, while liquid biopsy can provide noninvasive real-time monitoring via circulating biomarkers such as cell-free DNA, circulating tumor DNA, circulating tumor cells, and extracellular vesicles in plasma, cerebrospinal fluid, urine, and saliva; it advances beyond prior biomarker catalogs by delivering a quantitative technology scorecard comparing cfDNA-, CTC-, and EV-based platforms in terms of sensitivity, clinical actionability, cost, and scalability, and introduces a multimodal decision matrix and an AI-driven fusion pipeline integrating fragmentomics, EV proteomics, and CTC transcriptomics to enhance minimal residual disease detection a
This review focuses on liquid biopsy for glioblastoma, noting that tissue biopsy is invasive and fails to capture tumor heterogeneity or temporal dynamics, while liquid biopsy can provide noninvasive real-time monitoring via circulating biomarkers such as cell-free DNA, circulating tumor DNA, circulating tumor cells, and extracellular vesicles in plasma, cerebrospinal fluid, urine, and saliva; it advances beyond prior biomarker catalogs by delivering a quantitative technology scorecard comparing cfDNA-, CTC-, and EV-based platforms in terms of sensitivity, clinical actionability, cost, and scalability, and introduces a multimodal decision matrix and an AI-driven fusion pipeline integrating fragmentomics, EV proteomics, and CTC transcriptomics to enhance minimal residual disease detection a
This review focuses on liquid biopsy for glioblastoma, noting that tissue biopsy is invasive and fails to capture tumor heterogeneity or temporal dynamics, while liquid biopsy can provide noninvasive real-time monitoring via circulating biomarkers such as cell-free DNA, circulating tumor DNA, circulating tumor cells, and extracellular vesicles in plasma, cerebrospinal fluid, urine, and saliva; it advances beyond prior biomarker catalogs by delivering a quantitative technology scorecard comparing cfDNA-, CTC-, and EV-based platforms in terms of sensitivity, clinical actionability, cost, and scalability, and introduces a multimodal decision matrix and an AI-driven fusion pipeline integrating fragmentomics, EV proteomics, and CTC transcriptomics to enhance minimal residual disease detection a
medRxiv This systematic review and meta-analysis of 445 studies found that weaker preoperative social connection was associated with early postoperative mortality (OR 1.50, 95% CI 1.12-2.01) and non-home discharge (OR 1.95, 95% CI 1.35-2.81), while confidence intervals for postoperative survival, unplanned readmission, and complications all included 1, with overall certainty ranging from low to very low.
This systematic review and meta-analysis of 445 studies found that weaker preoperative social connection was associated with early postoperative mortality (OR 1.50, 95% CI 1.12-2.01) and non-home discharge (OR 1.95, 95% CI 1.35-2.81), while confidence intervals for postoperative survival, unplanned readmission, and complications all included 1, with overall certainty ranging from low to very low.
This systematic review and meta-analysis of 445 studies found that weaker preoperative social connection was associated with early postoperative mortality (OR 1.50, 95% CI 1.12-2.01) and non-home discharge (OR 1.95, 95% CI 1.35-2.81), while confidence intervals for postoperative survival, unplanned readmission, and complications all included 1, with overall certainty ranging from low to very low.
This systematic review and meta-analysis of 445 studies found that weaker preoperative social connection was associated with early postoperative mortality (OR 1.50, 95% CI 1.12-2.01) and non-home discharge (OR 1.95, 95% CI 1.35-2.81), while confidence intervals for postoperative survival, unplanned readmission, and complications all included 1, with overall certainty ranging from low to very low.
arXiv This work fine-tunes baseline models on 27 widely used benchmark datasets, categorizes heterophilic datasets into malignant, benign, and ambiguous groups based on whether graph-aware models underperform their coupled graph-agnostic counterparts, re-evaluates 11 SOTA heterophily-specific models, and conducts the first quantitative evaluation of 11 homophily metrics on synthetic graphs from three generation methods, finding that most SOTA models do not significantly outperform the best baselines and that classic metrics remain competitive.
This work fine-tunes baseline models on 27 widely used benchmark datasets, categorizes heterophilic datasets into malignant, benign, and ambiguous groups based on whether graph-aware models underperform their coupled graph-agnostic counterparts, re-evaluates 11 SOTA heterophily-specific models, and conducts the first quantitative evaluation of 11 homophily metrics on synthetic graphs from three generation methods, finding that most SOTA models do not significantly outperform the best baselines and that classic metrics remain competitive.
This work fine-tunes baseline models on 27 widely used benchmark datasets, categorizes heterophilic datasets into malignant, benign, and ambiguous groups based on whether graph-aware models underperform their coupled graph-agnostic counterparts, re-evaluates 11 SOTA heterophily-specific models, and conducts the first quantitative evaluation of 11 homophily metrics on synthetic graphs from three generation methods, finding that most SOTA models do not significantly outperform the best baselines and that classic metrics remain competitive.
This work fine-tunes baseline models on 27 widely used benchmark datasets, categorizes heterophilic datasets into malignant, benign, and ambiguous groups based on whether graph-aware models underperform their coupled graph-agnostic counterparts, re-evaluates 11 SOTA heterophily-specific models, and conducts the first quantitative evaluation of 11 homophily metrics on synthetic graphs from three generation methods, finding that most SOTA models do not significantly outperform the best baselines and that classic metrics remain competitive.
Biochimie This review synthesizes the 15 MAPK homologues across pathogenic Leishmania species, linking individual kinases to parasite differentiation, intracellular survival, stress tolerance, motility, virulence and drug response, distinguishing experimentally validated mechanistic targets from computationally proposed candidates, and evaluating small-molecule and natural-product inhibitors, drug repurposing, structure-guided approaches, and emerging contributions from molecular dynamics, AlphaFold-based modelling, artificial intelligence and nanotechnology, while noting evidence for MAPK10 as a vaccine-associated antigen.
This review synthesizes the 15 MAPK homologues across pathogenic Leishmania species, linking individual kinases to parasite differentiation, intracellular survival, stress tolerance, motility, virulence and drug response, distinguishing experimentally validated mechanistic targets from computationally proposed candidates, and evaluating small-molecule and natural-product inhibitors, drug repurposing, structure-guided approaches, and emerging contributions from molecular dynamics, AlphaFold-based modelling, artificial intelligence and nanotechnology, while noting evidence for MAPK10 as a vaccine-associated antigen.
This review synthesizes the 15 MAPK homologues across pathogenic Leishmania species, linking individual kinases to parasite differentiation, intracellular survival, stress tolerance, motility, virulence and drug response, distinguishing experimentally validated mechanistic targets from computationally proposed candidates, and evaluating small-molecule and natural-product inhibitors, drug repurposing, structure-guided approaches, and emerging contributions from molecular dynamics, AlphaFold-based modelling, artificial intelligence and nanotechnology, while noting evidence for MAPK10 as a vaccine-associated antigen.
This review synthesizes the 15 MAPK homologues across pathogenic Leishmania species, linking individual kinases to parasite differentiation, intracellular survival, stress tolerance, motility, virulence and drug response, distinguishing experimentally validated mechanistic targets from computationally proposed candidates, and evaluating small-molecule and natural-product inhibitors, drug repurposing, structure-guided approaches, and emerging contributions from molecular dynamics, AlphaFold-based modelling, artificial intelligence and nanotechnology, while noting evidence for MAPK10 as a vaccine-associated antigen.
JMIR AI This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.
This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.
This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.
This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.
Advanced science (Weinheim, Baden-Wurttemberg, Germany) This study applied AI-guided phenotypic drug repurposing to drug-resistant Streptococcus pneumoniae, using ensembles of transformer, graph, and tree models trained on 1849 actives and 34 503 inactives to prospectively examine 6747 drugs, selecting 11 candidate antibiotics of which nine strongly reduced in vitro growth of S. pneumoniae R6 (IC50 ≤ 0.4 µg/mL), with the most potent drugs thiostrepton and ceftiofur showing IC50 values of 0.0001 µg/mL (60.1 pM) and 0.0004 µg/mL (764 pM), respectively, and thiostrepton remaining highly potent against multidrug-resistant strains.
This study applied AI-guided phenotypic drug repurposing to drug-resistant Streptococcus pneumoniae, using ensembles of transformer, graph, and tree models trained on 1849 actives and 34 503 inactives to prospectively examine 6747 drugs, selecting 11 candidate antibiotics of which nine strongly reduced in vitro growth of S. pneumoniae R6 (IC50 ≤ 0.4 µg/mL), with the most potent drugs thiostrepton and ceftiofur showing IC50 values of 0.0001 µg/mL (60.1 pM) and 0.0004 µg/mL (764 pM), respectively, and thiostrepton remaining highly potent against multidrug-resistant strains.
This study applied AI-guided phenotypic drug repurposing to drug-resistant Streptococcus pneumoniae, using ensembles of transformer, graph, and tree models trained on 1849 actives and 34 503 inactives to prospectively examine 6747 drugs, selecting 11 candidate antibiotics of which nine strongly reduced in vitro growth of S. pneumoniae R6 (IC50 ≤ 0.4 µg/mL), with the most potent drugs thiostrepton and ceftiofur showing IC50 values of 0.0001 µg/mL (60.1 pM) and 0.0004 µg/mL (764 pM), respectively, and thiostrepton remaining highly potent against multidrug-resistant strains.
This study applied AI-guided phenotypic drug repurposing to drug-resistant Streptococcus pneumoniae, using ensembles of transformer, graph, and tree models trained on 1849 actives and 34 503 inactives to prospectively examine 6747 drugs, selecting 11 candidate antibiotics of which nine strongly reduced in vitro growth of S. pneumoniae R6 (IC50 ≤ 0.4 µg/mL), with the most potent drugs thiostrepton and ceftiofur showing IC50 values of 0.0001 µg/mL (60.1 pM) and 0.0004 µg/mL (764 pM), respectively, and thiostrepton remaining highly potent against multidrug-resistant strains.
Structure (London, England : 1993) This work presents AlphaBridge, a reproducible, objective, and automated toolkit, also available as a web server, that combines AlphaFold3 confidence metrics to cluster sequence motifs participating in binary interactions and subsequently in 3D interfaces of complexes, visualizes interaction interfaces within confidence limits via chord diagrams, network graphs, and summary tables of predicted interfaces and intermolecular interactions linked to interactive graphics, and was validated for scoring binary and multi-component protein complexes with real-life examples discussed.
This work presents AlphaBridge, a reproducible, objective, and automated toolkit, also available as a web server, that combines AlphaFold3 confidence metrics to cluster sequence motifs participating in binary interactions and subsequently in 3D interfaces of complexes, visualizes interaction interfaces within confidence limits via chord diagrams, network graphs, and summary tables of predicted interfaces and intermolecular interactions linked to interactive graphics, and was validated for scoring binary and multi-component protein complexes with real-life examples discussed.
This work presents AlphaBridge, a reproducible, objective, and automated toolkit, also available as a web server, that combines AlphaFold3 confidence metrics to cluster sequence motifs participating in binary interactions and subsequently in 3D interfaces of complexes, visualizes interaction interfaces within confidence limits via chord diagrams, network graphs, and summary tables of predicted interfaces and intermolecular interactions linked to interactive graphics, and was validated for scoring binary and multi-component protein complexes with real-life examples discussed.
This work presents AlphaBridge, a reproducible, objective, and automated toolkit, also available as a web server, that combines AlphaFold3 confidence metrics to cluster sequence motifs participating in binary interactions and subsequently in 3D interfaces of complexes, visualizes interaction interfaces within confidence limits via chord diagrams, network graphs, and summary tables of predicted interfaces and intermolecular interactions linked to interactive graphics, and was validated for scoring binary and multi-component protein complexes with real-life examples discussed.
The Permanente journal This article synthesizes early cross-project insights from the 5 projects funded by the Augmented Intelligence in Medicine and Healthcare Initiative (AIM-HI), led by Kaiser Permanente and funded by the Gordon and Betty Moore Foundation and selected through a national, multistage review process using a structured scoring rubric, covering sepsis management, venous thromboembolism risk assessment, diabetic retinopathy screening, cardiac amyloidosis detection, and pediatric asthma risk prediction; it reports that real-world AI deployment was feasible across varied clinical environments, that common challenges included electronic health record integration, data complexity, regulatory requirements, and variation in clinical workflows, and that implementation success depends on thoughtful integra
This article synthesizes early cross-project insights from the 5 projects funded by the Augmented Intelligence in Medicine and Healthcare Initiative (AIM-HI), led by Kaiser Permanente and funded by the Gordon and Betty Moore Foundation and selected through a national, multistage review process using a structured scoring rubric, covering sepsis management, venous thromboembolism risk assessment, diabetic retinopathy screening, cardiac amyloidosis detection, and pediatric asthma risk prediction; it reports that real-world AI deployment was feasible across varied clinical environments, that common challenges included electronic health record integration, data complexity, regulatory requirements, and variation in clinical workflows, and that implementation success depends on thoughtful integra
This article synthesizes early cross-project insights from the 5 projects funded by the Augmented Intelligence in Medicine and Healthcare Initiative (AIM-HI), led by Kaiser Permanente and funded by the Gordon and Betty Moore Foundation and selected through a national, multistage review process using a structured scoring rubric, covering sepsis management, venous thromboembolism risk assessment, diabetic retinopathy screening, cardiac amyloidosis detection, and pediatric asthma risk prediction; it reports that real-world AI deployment was feasible across varied clinical environments, that common challenges included electronic health record integration, data complexity, regulatory requirements, and variation in clinical workflows, and that implementation success depends on thoughtful integra
This article synthesizes early cross-project insights from the 5 projects funded by the Augmented Intelligence in Medicine and Healthcare Initiative (AIM-HI), led by Kaiser Permanente and funded by the Gordon and Betty Moore Foundation and selected through a national, multistage review process using a structured scoring rubric, covering sepsis management, venous thromboembolism risk assessment, diabetic retinopathy screening, cardiac amyloidosis detection, and pediatric asthma risk prediction; it reports that real-world AI deployment was feasible across varied clinical environments, that common challenges included electronic health record integration, data complexity, regulatory requirements, and variation in clinical workflows, and that implementation success depends on thoughtful integra
medRxiv Using AlphaGenome and AlphaMissense to quantify the disruption imposed by somatic mutations across 8,800 patients and 33 cancer types from The Cancer Genome Atlas, this study found that recurrent hotspot mutations showed substantially larger predicted protein-level effects while non-hotspot mutations exhibited larger regulatory effects across most cancer types; aggregating variant-level predictions into patient-gene disruption profiles capturing transcriptional activity, chromatin accessibility, transcription factor binding, and splicing yielded gene- and modality-specific profiles that reflected tissue of origin, cancer type, and microsatellite-instability status while retaining information beyond tumor mutational burden; among patients lacking recurrent hotspot mutations, higher predicte
Using AlphaGenome and AlphaMissense to quantify the disruption imposed by somatic mutations across 8,800 patients and 33 cancer types from The Cancer Genome Atlas, this study found that recurrent hotspot mutations showed substantially larger predicted protein-level effects while non-hotspot mutations exhibited larger regulatory effects across most cancer types; aggregating variant-level predictions into patient-gene disruption profiles capturing transcriptional activity, chromatin accessibility, transcription factor binding, and splicing yielded gene- and modality-specific profiles that reflected tissue of origin, cancer type, and microsatellite-instability status while retaining information beyond tumor mutational burden; among patients lacking recurrent hotspot mutations, higher predicte
Using AlphaGenome and AlphaMissense to quantify the disruption imposed by somatic mutations across 8,800 patients and 33 cancer types from The Cancer Genome Atlas, this study found that recurrent hotspot mutations showed substantially larger predicted protein-level effects while non-hotspot mutations exhibited larger regulatory effects across most cancer types; aggregating variant-level predictions into patient-gene disruption profiles capturing transcriptional activity, chromatin accessibility, transcription factor binding, and splicing yielded gene- and modality-specific profiles that reflected tissue of origin, cancer type, and microsatellite-instability status while retaining information beyond tumor mutational burden; among patients lacking recurrent hotspot mutations, higher predicte
Using AlphaGenome and AlphaMissense to quantify the disruption imposed by somatic mutations across 8,800 patients and 33 cancer types from The Cancer Genome Atlas, this study found that recurrent hotspot mutations showed substantially larger predicted protein-level effects while non-hotspot mutations exhibited larger regulatory effects across most cancer types; aggregating variant-level predictions into patient-gene disruption profiles capturing transcriptional activity, chromatin accessibility, transcription factor binding, and splicing yielded gene- and modality-specific profiles that reflected tissue of origin, cancer type, and microsatellite-instability status while retaining information beyond tumor mutational burden; among patients lacking recurrent hotspot mutations, higher predicte
Practical radiation oncology This study implemented AP-FLART, which integrates dosimetric score-based beam angle selection, multi-modality-guided dose prediction, and function-guided dose mimicking, within RayStation and evaluated it on a test dataset of 33 lung cancer patients who underwent SPECT ventilation or perfusion imaging and lung radiotherapy, finding that automatic FLART plans significantly reduced high-function lung mean dose by 15.1% versus manual conventional radiotherapy plans, lowered the probability of grade >=2 radiation pneumonitis by 6.25 percentage points (27%) among FLART-benefiting patients, achieved clinical benefits similar to manual FLART plans, were clinically acceptable without modification in 87.9% of cases, and cut planning time from 2-3 hours to approximately 8 minutes.
This study implemented AP-FLART, which integrates dosimetric score-based beam angle selection, multi-modality-guided dose prediction, and function-guided dose mimicking, within RayStation and evaluated it on a test dataset of 33 lung cancer patients who underwent SPECT ventilation or perfusion imaging and lung radiotherapy, finding that automatic FLART plans significantly reduced high-function lung mean dose by 15.1% versus manual conventional radiotherapy plans, lowered the probability of grade >=2 radiation pneumonitis by 6.25 percentage points (27%) among FLART-benefiting patients, achieved clinical benefits similar to manual FLART plans, were clinically acceptable without modification in 87.9% of cases, and cut planning time from 2-3 hours to approximately 8 minutes.
This study implemented AP-FLART, which integrates dosimetric score-based beam angle selection, multi-modality-guided dose prediction, and function-guided dose mimicking, within RayStation and evaluated it on a test dataset of 33 lung cancer patients who underwent SPECT ventilation or perfusion imaging and lung radiotherapy, finding that automatic FLART plans significantly reduced high-function lung mean dose by 15.1% versus manual conventional radiotherapy plans, lowered the probability of grade >=2 radiation pneumonitis by 6.25 percentage points (27%) among FLART-benefiting patients, achieved clinical benefits similar to manual FLART plans, were clinically acceptable without modification in 87.9% of cases, and cut planning time from 2-3 hours to approximately 8 minutes.
This study implemented AP-FLART, which integrates dosimetric score-based beam angle selection, multi-modality-guided dose prediction, and function-guided dose mimicking, within RayStation and evaluated it on a test dataset of 33 lung cancer patients who underwent SPECT ventilation or perfusion imaging and lung radiotherapy, finding that automatic FLART plans significantly reduced high-function lung mean dose by 15.1% versus manual conventional radiotherapy plans, lowered the probability of grade >=2 radiation pneumonitis by 6.25 percentage points (27%) among FLART-benefiting patients, achieved clinical benefits similar to manual FLART plans, were clinically acceptable without modification in 87.9% of cases, and cut planning time from 2-3 hours to approximately 8 minutes.
Journal of Nondestructive Evaluation This study conditioned concrete prisms of four compressive strengths (f′c = 32–56 MPa) and two alkali-silica reaction (ASR) damaged specimens by compressive loading, monitored the subsequent velocity recovery with coda wave interferometry (CWI), and applied a self-referencing temperature correction to remove drift from ambient fluctuations of about ±0.2°C; for intact specimens the recovery rate mv and velocity drop magnitude |c| increased with strength (|c| from 3.08×10⁻⁴ to 6.02×10⁻⁴, mv from 6.77×10⁻⁵/s to 13.76×10⁻⁵/s) and recovery time shortened from 9.9 h to 6.
This study conditioned concrete prisms of four compressive strengths (f′c = 32–56 MPa) and two alkali-silica reaction (ASR) damaged specimens by compressive loading, monitored the subsequent velocity recovery with coda wave interferometry (CWI), and applied a self-referencing temperature correction to remove drift from ambient fluctuations of about ±0.2°C; for intact specimens the recovery rate mv and velocity drop magnitude |c| increased with strength (|c| from 3.08×10⁻⁴ to 6.02×10⁻⁴, mv from 6.77×10⁻⁵/s to 13.76×10⁻⁵/s) and recovery time shortened from 9.9 h to 6.
This study conditioned concrete prisms of four compressive strengths (f′c = 32–56 MPa) and two alkali-silica reaction (ASR) damaged specimens by compressive loading, monitored the subsequent velocity recovery with coda wave interferometry (CWI), and applied a self-referencing temperature correction to remove drift from ambient fluctuations of about ±0.2°C; for intact specimens the recovery rate mv and velocity drop magnitude |c| increased with strength (|c| from 3.08×10⁻⁴ to 6.02×10⁻⁴, mv from 6.77×10⁻⁵/s to 13.76×10⁻⁵/s) and recovery time shortened from 9.9 h to 6.
This study conditioned concrete prisms of four compressive strengths (f′c = 32–56 MPa) and two alkali-silica reaction (ASR) damaged specimens by compressive loading, monitored the subsequent velocity recovery with coda wave interferometry (CWI), and applied a self-referencing temperature correction to remove drift from ambient fluctuations of about ±0.2°C; for intact specimens the recovery rate mv and velocity drop magnitude |c| increased with strength (|c| from 3.08×10⁻⁴ to 6.02×10⁻⁴, mv from 6.77×10⁻⁵/s to 13.76×10⁻⁵/s) and recovery time shortened from 9.9 h to 6.
Research Square Using 903 open-ended responses across six variables from a European PhD student survey, this study had five human coders and GPT-5.4 each perform the same inductive content analysis procedure to produce codes and themes, and measured agreement with the Adjusted Rand Index (ARI), finding average human-LLM agreement of 0.61 for coding and 0.54 for themes, close to within-human consistency (0.68) and within-LLM consistency (0.76), with wide variation across variables and low within-entity consistency consistently accompanying low between-entity agreement.
Using 903 open-ended responses across six variables from a European PhD student survey, this study had five human coders and GPT-5.4 each perform the same inductive content analysis procedure to produce codes and themes, and measured agreement with the Adjusted Rand Index (ARI), finding average human-LLM agreement of 0.61 for coding and 0.54 for themes, close to within-human consistency (0.68) and within-LLM consistency (0.76), with wide variation across variables and low within-entity consistency consistently accompanying low between-entity agreement.
Using 903 open-ended responses across six variables from a European PhD student survey, this study had five human coders and GPT-5.4 each perform the same inductive content analysis procedure to produce codes and themes, and measured agreement with the Adjusted Rand Index (ARI), finding average human-LLM agreement of 0.61 for coding and 0.54 for themes, close to within-human consistency (0.68) and within-LLM consistency (0.76), with wide variation across variables and low within-entity consistency consistently accompanying low between-entity agreement.
Using 903 open-ended responses across six variables from a European PhD student survey, this study had five human coders and GPT-5.4 each perform the same inductive content analysis procedure to produce codes and themes, and measured agreement with the Adjusted Rand Index (ARI), finding average human-LLM agreement of 0.61 for coding and 0.54 for themes, close to within-human consistency (0.68) and within-LLM consistency (0.76), with wide variation across variables and low within-entity consistency consistently accompanying low between-entity agreement.
RNA This work develops a continuum-based reaction-diffusion model of XIST-mediated gene silencing spread on chromosomes, finding that XIST spread can be tuned by known negative feedback loops regulating its synthesis and degradation, while silencing spread is controlled by a wave-pinning mechanism driven by global regulation of the silencing complex together with local epigenetic regulators, and uses a 3D chromosome structure inferred from experimental data to show spatiotemporal regulation of silencing spread.
This work develops a continuum-based reaction-diffusion model of XIST-mediated gene silencing spread on chromosomes, finding that XIST spread can be tuned by known negative feedback loops regulating its synthesis and degradation, while silencing spread is controlled by a wave-pinning mechanism driven by global regulation of the silencing complex together with local epigenetic regulators, and uses a 3D chromosome structure inferred from experimental data to show spatiotemporal regulation of silencing spread.
This work develops a continuum-based reaction-diffusion model of XIST-mediated gene silencing spread on chromosomes, finding that XIST spread can be tuned by known negative feedback loops regulating its synthesis and degradation, while silencing spread is controlled by a wave-pinning mechanism driven by global regulation of the silencing complex together with local epigenetic regulators, and uses a 3D chromosome structure inferred from experimental data to show spatiotemporal regulation of silencing spread.
This work develops a continuum-based reaction-diffusion model of XIST-mediated gene silencing spread on chromosomes, finding that XIST spread can be tuned by known negative feedback loops regulating its synthesis and degradation, while silencing spread is controlled by a wave-pinning mechanism driven by global regulation of the silencing complex together with local epigenetic regulators, and uses a 3D chromosome structure inferred from experimental data to show spatiotemporal regulation of silencing spread.
npj Digital Medicine The study introduces a domain-agnostic paired-comparison approach that fine-tunes a large language model to emulate documented emergency triage decisions and then compares predictions on sex-swapped pairs in which only sex is flipped while documented clinical content is held constant, finding that otherwise identical presentations were more likely to receive a lower-severity (less urgent) predicted triage score as female than male, with a pooled per-pair rate of about 1.1% (95% CI 0.9–1.3) across more than 140,000 Bordeaux University Hospital admissions and a directionally consistent but larger 2.2% (1.7–2.7) in MIMIC-IV, while a model retrained on sex-neutralized inputs eliminated the between-sex prediction gap, indicating the asymmetry is mediated by explicit sex markers.
The study introduces a domain-agnostic paired-comparison approach that fine-tunes a large language model to emulate documented emergency triage decisions and then compares predictions on sex-swapped pairs in which only sex is flipped while documented clinical content is held constant, finding that otherwise identical presentations were more likely to receive a lower-severity (less urgent) predicted triage score as female than male, with a pooled per-pair rate of about 1.1% (95% CI 0.9–1.3) across more than 140,000 Bordeaux University Hospital admissions and a directionally consistent but larger 2.2% (1.7–2.7) in MIMIC-IV, while a model retrained on sex-neutralized inputs eliminated the between-sex prediction gap, indicating the asymmetry is mediated by explicit sex markers.
The study introduces a domain-agnostic paired-comparison approach that fine-tunes a large language model to emulate documented emergency triage decisions and then compares predictions on sex-swapped pairs in which only sex is flipped while documented clinical content is held constant, finding that otherwise identical presentations were more likely to receive a lower-severity (less urgent) predicted triage score as female than male, with a pooled per-pair rate of about 1.1% (95% CI 0.9–1.3) across more than 140,000 Bordeaux University Hospital admissions and a directionally consistent but larger 2.2% (1.7–2.7) in MIMIC-IV, while a model retrained on sex-neutralized inputs eliminated the between-sex prediction gap, indicating the asymmetry is mediated by explicit sex markers.
The study introduces a domain-agnostic paired-comparison approach that fine-tunes a large language model to emulate documented emergency triage decisions and then compares predictions on sex-swapped pairs in which only sex is flipped while documented clinical content is held constant, finding that otherwise identical presentations were more likely to receive a lower-severity (less urgent) predicted triage score as female than male, with a pooled per-pair rate of about 1.1% (95% CI 0.9–1.3) across more than 140,000 Bordeaux University Hospital admissions and a directionally consistent but larger 2.2% (1.7–2.7) in MIMIC-IV, while a model retrained on sex-neutralized inputs eliminated the between-sex prediction gap, indicating the asymmetry is mediated by explicit sex markers.
Anesthesiology Using 203,434 ICU admissions from 2001 to 2022 across more than 200 hospitals in the MIMIC-III, MIMIC-IV, eICU, and HiRID databases, the study developed and externally validated a multimodal deep-learning model that predicts subsequent inpatient mortality from time-invariant variables, time-variant variables, clinical notes, and chest x-ray images available within the first 24 h of ICU admission; with structured data alone the model reached an AUROC of 0.92 (95% CI, 0.90 to 0.93), an AUPRC of 0.53 (95% CI, 0.49 to 0.57), and a Brier score of 0.19 (95% CI, 0.18 to 0.20), external validation across eight eICU institutions yielded AUROCs of 0.84 to 0.92, and in the subgroup with both notes and imaging, adding text and images raised the AUROC modestly from 0.87 (95% CI, 0.85 to 0.89) to 0.
Using 203,434 ICU admissions from 2001 to 2022 across more than 200 hospitals in the MIMIC-III, MIMIC-IV, eICU, and HiRID databases, the study developed and externally validated a multimodal deep-learning model that predicts subsequent inpatient mortality from time-invariant variables, time-variant variables, clinical notes, and chest x-ray images available within the first 24 h of ICU admission; with structured data alone the model reached an AUROC of 0.92 (95% CI, 0.90 to 0.93), an AUPRC of 0.53 (95% CI, 0.49 to 0.57), and a Brier score of 0.19 (95% CI, 0.18 to 0.20), external validation across eight eICU institutions yielded AUROCs of 0.84 to 0.92, and in the subgroup with both notes and imaging, adding text and images raised the AUROC modestly from 0.87 (95% CI, 0.85 to 0.89) to 0.
Using 203,434 ICU admissions from 2001 to 2022 across more than 200 hospitals in the MIMIC-III, MIMIC-IV, eICU, and HiRID databases, the study developed and externally validated a multimodal deep-learning model that predicts subsequent inpatient mortality from time-invariant variables, time-variant variables, clinical notes, and chest x-ray images available within the first 24 h of ICU admission; with structured data alone the model reached an AUROC of 0.92 (95% CI, 0.90 to 0.93), an AUPRC of 0.53 (95% CI, 0.49 to 0.57), and a Brier score of 0.19 (95% CI, 0.18 to 0.20), external validation across eight eICU institutions yielded AUROCs of 0.84 to 0.92, and in the subgroup with both notes and imaging, adding text and images raised the AUROC modestly from 0.87 (95% CI, 0.85 to 0.89) to 0.
Using 203,434 ICU admissions from 2001 to 2022 across more than 200 hospitals in the MIMIC-III, MIMIC-IV, eICU, and HiRID databases, the study developed and externally validated a multimodal deep-learning model that predicts subsequent inpatient mortality from time-invariant variables, time-variant variables, clinical notes, and chest x-ray images available within the first 24 h of ICU admission; with structured data alone the model reached an AUROC of 0.92 (95% CI, 0.90 to 0.93), an AUPRC of 0.53 (95% CI, 0.49 to 0.57), and a Brier score of 0.19 (95% CI, 0.18 to 0.20), external validation across eight eICU institutions yielded AUROCs of 0.84 to 0.92, and in the subgroup with both notes and imaging, adding text and images raised the AUROC modestly from 0.87 (95% CI, 0.85 to 0.89) to 0.
medRxiv This study updated an automated tumour infiltrating lymphocyte (TIL) assessment pipeline, proposing TRIGS aligned with clinical scoring guidelines and foundation model-based SAM-TIL, and showed in 166 neoadjuvantly treated patients in the TransNEO cohort that they predicted pathological complete response with odds ratios of 1.95 (95% CI 1.22-3.03, p=0.005) and 2.32 (95% CI 1.43-3.77, p=0.001), and in 277 triple negative and HER2-positive patients in The Cancer Genome Atlas showed overall survival hazard ratios of 0.79 (95% CI 0.63-1.00, p=0.05) and 0.80 (95% CI 0.67-0.97, p=0.02), with correlation to gold standard clinical assessment of 0.59-0.69 and no substantial difference from gold standard assessment in predicting pathological complete response (AUC 0.60-0.
This study updated an automated tumour infiltrating lymphocyte (TIL) assessment pipeline, proposing TRIGS aligned with clinical scoring guidelines and foundation model-based SAM-TIL, and showed in 166 neoadjuvantly treated patients in the TransNEO cohort that they predicted pathological complete response with odds ratios of 1.95 (95% CI 1.22-3.03, p=0.005) and 2.32 (95% CI 1.43-3.77, p=0.001), and in 277 triple negative and HER2-positive patients in The Cancer Genome Atlas showed overall survival hazard ratios of 0.79 (95% CI 0.63-1.00, p=0.05) and 0.80 (95% CI 0.67-0.97, p=0.02), with correlation to gold standard clinical assessment of 0.59-0.69 and no substantial difference from gold standard assessment in predicting pathological complete response (AUC 0.60-0.
This study updated an automated tumour infiltrating lymphocyte (TIL) assessment pipeline, proposing TRIGS aligned with clinical scoring guidelines and foundation model-based SAM-TIL, and showed in 166 neoadjuvantly treated patients in the TransNEO cohort that they predicted pathological complete response with odds ratios of 1.95 (95% CI 1.22-3.03, p=0.005) and 2.32 (95% CI 1.43-3.77, p=0.001), and in 277 triple negative and HER2-positive patients in The Cancer Genome Atlas showed overall survival hazard ratios of 0.79 (95% CI 0.63-1.00, p=0.05) and 0.80 (95% CI 0.67-0.97, p=0.02), with correlation to gold standard clinical assessment of 0.59-0.69 and no substantial difference from gold standard assessment in predicting pathological complete response (AUC 0.60-0.
This study updated an automated tumour infiltrating lymphocyte (TIL) assessment pipeline, proposing TRIGS aligned with clinical scoring guidelines and foundation model-based SAM-TIL, and showed in 166 neoadjuvantly treated patients in the TransNEO cohort that they predicted pathological complete response with odds ratios of 1.95 (95% CI 1.22-3.03, p=0.005) and 2.32 (95% CI 1.43-3.77, p=0.001), and in 277 triple negative and HER2-positive patients in The Cancer Genome Atlas showed overall survival hazard ratios of 0.79 (95% CI 0.63-1.00, p=0.05) and 0.80 (95% CI 0.67-0.97, p=0.02), with correlation to gold standard clinical assessment of 0.59-0.69 and no substantial difference from gold standard assessment in predicting pathological complete response (AUC 0.60-0.
The latest research from Google The work proposes Retrieve-for-Train, which first trains a fan-out language model with offline reinforcement learning (built on Gemma3-4B and Qwen3-4B, emitting 10 sub-queries per prompt) under a composite reward of groundedness, Vendi-Score diversity, and alignment that scores the whole result set, then distills that behavior into a 53.9M-parameter diffusion retriever that generates the complete target set in one non-autoregressive parallel pass in continuous embedding space, outperforming single-query search, zero-shot expansion, and a Best-of-N baseline on open-ended abstract retrieval and weakly supervised compositional retrieval while achieving a 12 to 20 speedup over autoregressive approaches.
The work proposes Retrieve-for-Train, which first trains a fan-out language model with offline reinforcement learning (built on Gemma3-4B and Qwen3-4B, emitting 10 sub-queries per prompt) under a composite reward of groundedness, Vendi-Score diversity, and alignment that scores the whole result set, then distills that behavior into a 53.9M-parameter diffusion retriever that generates the complete target set in one non-autoregressive parallel pass in continuous embedding space, outperforming single-query search, zero-shot expansion, and a Best-of-N baseline on open-ended abstract retrieval and weakly supervised compositional retrieval while achieving a 12 to 20 speedup over autoregressive approaches.
The work proposes Retrieve-for-Train, which first trains a fan-out language model with offline reinforcement learning (built on Gemma3-4B and Qwen3-4B, emitting 10 sub-queries per prompt) under a composite reward of groundedness, Vendi-Score diversity, and alignment that scores the whole result set, then distills that behavior into a 53.9M-parameter diffusion retriever that generates the complete target set in one non-autoregressive parallel pass in continuous embedding space, outperforming single-query search, zero-shot expansion, and a Best-of-N baseline on open-ended abstract retrieval and weakly supervised compositional retrieval while achieving a 12 to 20 speedup over autoregressive approaches.
The work proposes Retrieve-for-Train, which first trains a fan-out language model with offline reinforcement learning (built on Gemma3-4B and Qwen3-4B, emitting 10 sub-queries per prompt) under a composite reward of groundedness, Vendi-Score diversity, and alignment that scores the whole result set, then distills that behavior into a 53.9M-parameter diffusion retriever that generates the complete target set in one non-autoregressive parallel pass in continuous embedding space, outperforming single-query search, zero-shot expansion, and a Best-of-N baseline on open-ended abstract retrieval and weakly supervised compositional retrieval while achieving a 12 to 20 speedup over autoregressive approaches.
Nature News According to a Nature news report, Paper2Agent reads a paper's main text, code and data sets, deposits them on an MCP server, and has a team of AI agents autonomously build tools that apply the paper's methods, producing a paper-specific agent that can be questioned in plain language as a 'virtual corresponding author'; the report says it created an agent for the AlphaGenome paper in about 45 minutes at US$14 of computing cost, that the agent answered genetics questions with near-perfect accuracy and outscored other top biomedical AI agents including Biomni, and that it was used to re-examine which causal gene explains a single-letter DNA change linked to 'bad' cholesterol, arriving at a different gene from the one pinpointed in the original paper.
According to a Nature news report, Paper2Agent reads a paper's main text, code and data sets, deposits them on an MCP server, and has a team of AI agents autonomously build tools that apply the paper's methods, producing a paper-specific agent that can be questioned in plain language as a 'virtual corresponding author'; the report says it created an agent for the AlphaGenome paper in about 45 minutes at US$14 of computing cost, that the agent answered genetics questions with near-perfect accuracy and outscored other top biomedical AI agents including Biomni, and that it was used to re-examine which causal gene explains a single-letter DNA change linked to 'bad' cholesterol, arriving at a different gene from the one pinpointed in the original paper.
According to a Nature news report, Paper2Agent reads a paper's main text, code and data sets, deposits them on an MCP server, and has a team of AI agents autonomously build tools that apply the paper's methods, producing a paper-specific agent that can be questioned in plain language as a 'virtual corresponding author'; the report says it created an agent for the AlphaGenome paper in about 45 minutes at US$14 of computing cost, that the agent answered genetics questions with near-perfect accuracy and outscored other top biomedical AI agents including Biomni, and that it was used to re-examine which causal gene explains a single-letter DNA change linked to 'bad' cholesterol, arriving at a different gene from the one pinpointed in the original paper.
According to a Nature news report, Paper2Agent reads a paper's main text, code and data sets, deposits them on an MCP server, and has a team of AI agents autonomously build tools that apply the paper's methods, producing a paper-specific agent that can be questioned in plain language as a 'virtual corresponding author'; the report says it created an agent for the AlphaGenome paper in about 45 minutes at US$14 of computing cost, that the agent answered genetics questions with near-perfect accuracy and outscored other top biomedical AI agents including Biomni, and that it was used to re-examine which causal gene explains a single-letter DNA change linked to 'bad' cholesterol, arriving at a different gene from the one pinpointed in the original paper.
Nature News By genetically engineering mice so that the precursor cells that would have formed the cerebral cortex do not survive, thereby emptying part of the brain cavity, and then implanting human brain tissue into newborn mice, the study reports that the graft expanded nearly fivefold within two to three months, filled more than 90% of the vacant space, sent projections deep into the rodents' spinal cords, and developed specialized neurons including cells similar to von Economo neurons, while behavioral tests showed no enhancement of the rodents' intellect.
By genetically engineering mice so that the precursor cells that would have formed the cerebral cortex do not survive, thereby emptying part of the brain cavity, and then implanting human brain tissue into newborn mice, the study reports that the graft expanded nearly fivefold within two to three months, filled more than 90% of the vacant space, sent projections deep into the rodents' spinal cords, and developed specialized neurons including cells similar to von Economo neurons, while behavioral tests showed no enhancement of the rodents' intellect.
By genetically engineering mice so that the precursor cells that would have formed the cerebral cortex do not survive, thereby emptying part of the brain cavity, and then implanting human brain tissue into newborn mice, the study reports that the graft expanded nearly fivefold within two to three months, filled more than 90% of the vacant space, sent projections deep into the rodents' spinal cords, and developed specialized neurons including cells similar to von Economo neurons, while behavioral tests showed no enhancement of the rodents' intellect.
By genetically engineering mice so that the precursor cells that would have formed the cerebral cortex do not survive, thereby emptying part of the brain cavity, and then implanting human brain tissue into newborn mice, the study reports that the graft expanded nearly fivefold within two to three months, filled more than 90% of the vacant space, sent projections deep into the rodents' spinal cords, and developed specialized neurons including cells similar to von Economo neurons, while behavioral tests showed no enhancement of the rodents' intellect.
Nature News This evidence bundle, comprising a Nature news feature and two bioRxiv preprints, reports progress in culturing Asgard archaea (Promethearchaeota): live-cell microscopy showed that cells of the Loki and Hodarchaea lineages drastically change shape on a minute timescale, extend and retract protrusions at 1.5 to 5.3 micrometres per minute, and crawl on glass surfaces, with actin inhibitors arresting these dynamics; separately, an Asgard archaeal virus infecting a novel strain of Ca. Lokiarchaeum ossiferum B36 was cultured for the first time, with a 16 kbp integrated provirus able to excise and replicate independently to form virus particles, placed in a new family Fylgjaviridae, while the host carries Septu, Wadjet and type II CBASS antiviral defence systems.
This evidence bundle, comprising a Nature news feature and two bioRxiv preprints, reports progress in culturing Asgard archaea (Promethearchaeota): live-cell microscopy showed that cells of the Loki and Hodarchaea lineages drastically change shape on a minute timescale, extend and retract protrusions at 1.5 to 5.3 micrometres per minute, and crawl on glass surfaces, with actin inhibitors arresting these dynamics; separately, an Asgard archaeal virus infecting a novel strain of Ca. Lokiarchaeum ossiferum B36 was cultured for the first time, with a 16 kbp integrated provirus able to excise and replicate independently to form virus particles, placed in a new family Fylgjaviridae, while the host carries Septu, Wadjet and type II CBASS antiviral defence systems.
This evidence bundle, comprising a Nature news feature and two bioRxiv preprints, reports progress in culturing Asgard archaea (Promethearchaeota): live-cell microscopy showed that cells of the Loki and Hodarchaea lineages drastically change shape on a minute timescale, extend and retract protrusions at 1.5 to 5.3 micrometres per minute, and crawl on glass surfaces, with actin inhibitors arresting these dynamics; separately, an Asgard archaeal virus infecting a novel strain of Ca. Lokiarchaeum ossiferum B36 was cultured for the first time, with a 16 kbp integrated provirus able to excise and replicate independently to form virus particles, placed in a new family Fylgjaviridae, while the host carries Septu, Wadjet and type II CBASS antiviral defence systems.
This evidence bundle, comprising a Nature news feature and two bioRxiv preprints, reports progress in culturing Asgard archaea (Promethearchaeota): live-cell microscopy showed that cells of the Loki and Hodarchaea lineages drastically change shape on a minute timescale, extend and retract protrusions at 1.5 to 5.3 micrometres per minute, and crawl on glass surfaces, with actin inhibitors arresting these dynamics; separately, an Asgard archaeal virus infecting a novel strain of Ca. Lokiarchaeum ossiferum B36 was cultured for the first time, with a 16 kbp integrated provirus able to excise and replicate independently to form virus particles, placed in a new family Fylgjaviridae, while the host carries Septu, Wadjet and type II CBASS antiviral defence systems.
Nature News Combining mass spectrometry, RNA sequencing and ribosomal profiling on 608 post-mortem dorsolateral prefrontal cortex samples from individuals with and without Alzheimer's disease, the study identified 4,321 microproteins (3,217 of them not previously characterized in the UniProtKB/Swiss-Prot standard human protein catalogue), applied a deep-learning model to rank the mass-spectrometry identification confidence of 3,001 of them with 1,067 rated high confidence, and found dozens of microproteins with altered expression in people with Alzheimer's disease, producing what is described as the largest atlas of microproteins in Alzheimer's disease made so far and making the dataset publicly available.
Combining mass spectrometry, RNA sequencing and ribosomal profiling on 608 post-mortem dorsolateral prefrontal cortex samples from individuals with and without Alzheimer's disease, the study identified 4,321 microproteins (3,217 of them not previously characterized in the UniProtKB/Swiss-Prot standard human protein catalogue), applied a deep-learning model to rank the mass-spectrometry identification confidence of 3,001 of them with 1,067 rated high confidence, and found dozens of microproteins with altered expression in people with Alzheimer's disease, producing what is described as the largest atlas of microproteins in Alzheimer's disease made so far and making the dataset publicly available.
Combining mass spectrometry, RNA sequencing and ribosomal profiling on 608 post-mortem dorsolateral prefrontal cortex samples from individuals with and without Alzheimer's disease, the study identified 4,321 microproteins (3,217 of them not previously characterized in the UniProtKB/Swiss-Prot standard human protein catalogue), applied a deep-learning model to rank the mass-spectrometry identification confidence of 3,001 of them with 1,067 rated high confidence, and found dozens of microproteins with altered expression in people with Alzheimer's disease, producing what is described as the largest atlas of microproteins in Alzheimer's disease made so far and making the dataset publicly available.
Combining mass spectrometry, RNA sequencing and ribosomal profiling on 608 post-mortem dorsolateral prefrontal cortex samples from individuals with and without Alzheimer's disease, the study identified 4,321 microproteins (3,217 of them not previously characterized in the UniProtKB/Swiss-Prot standard human protein catalogue), applied a deep-learning model to rank the mass-spectrometry identification confidence of 3,001 of them with 1,067 rated high confidence, and found dozens of microproteins with altered expression in people with Alzheimer's disease, producing what is described as the largest atlas of microproteins in Alzheimer's disease made so far and making the dataset publicly available.