Skip to main content

Daily report

AI and science frontiers · 2026-09-21

Only content delivered through the publication boundary on this date is included.

Microsoft Research

RetroChimera: Improving Small-Molecule Retrosynthesis Prediction with a Learned Two-Model Ensemble

This work presents RetroChimera, a retrosynthesis prediction framework that combines a Transformer-based de-novo model, R-SMILES 2, with a graph-neural-network model grounded in reaction templates, NeuralLoc, through a learned, rank-dependent ensembling strategy, so that their complementary strengths are leveraged to perform strongly across both common and rare reaction classes and, in blind tests, to produce disconnections of complex molecules that PhD-level chemists preferred over those from its constituent sub-models, from more established approaches, and even from the test set itself.
NVIDIA Research

AI Security Is an Engineering Problem: Solving It at Every Layer of the Agent Stack

This article argues that AI security should be treated as an engineering problem requiring defined security requirements, enforceable controls, named owners, and evidence that protections work, and proposes that security depends on the full agent stack (models, harnesses, runtime environments) working together, introducing NVIDIA OpenShell as an open source secure runtime with integrations across Open Secure AI Alliance partners, along with defensive tool examples such as CrowdStrike SafeMind, Palo Alto Networks Prisma AIRS, Capital One VulnHunter, and ReversingLabs Spectra Assure, while emphasizing open sharing of failure evidence and verification methods to shift advantage toward defenders.
Hugging Face

Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

This work reformulates block removal in large language models as a constrained binary optimization equivalent to finding low-energy states of an Ising glass, using a once-computed approximate-Hessian energy as a cheap proxy for downstream quality so that vast numbers of candidate configurations can be ranked without benchmarking any of them; on Llama-3.3-70B-Instruct at 50% compression (40 of 80 blocks removed) without retraining, MMLU stays near 77 while the strongest baseline, block influence, falls to about 54, with transfer also shown on Llama-3.1-8B-Instruct, Qwen3-14B, and the hybrid NVIDIA-Nemotron-3-Nano-30B-A3B-FP8.
OpenAI

Higgsfield AI ships new video features in a day with GPT-6 Astra

This report describes how Higgsfield AI uses GPT-6 Astra to let small businesses generate video ad variations from a single prompt and to let one engineer deliver new exploration features within a day, with the company attributing that speed to Astra's long-horizon task planning and close collaboration between its creative team and engineers.
OpenAI

Building Standards for the Next Phase of AI: A Policy Proposal for International Technical Standards on Frontier AI

This is a policy position paper arguing that the United States should lead an effort with countries worldwide to develop global technical standards for frontier AI—especially automated AI research and recursive self-improvement (RSI)—to address fragmentation, collective-action problems, and uneven capacity, while specifying that such standards should focus on capability measurement, risk assessment, and safeguard sufficiency rather than licenses or mandatory pre-release approval.
deepcybo-physai.github.io

PhysBrain 1.5: A Physical Foundation Model Unifying Language, Action, and Future World States in One Autoregressive Vocabulary

PhysBrain 1.5 builds on a pretrained Qwen3-VL backbone extended with dedicated action and visual-state tokens, formulating language responses, structured spatial outputs, end-effector trajectories, and future world states as discrete tokens jointly learned under a single next-token prediction objective, with embodied pretraining supervision drawn entirely from human interaction videos (egocentric, synchronized ego-exocentric, and panoramic recordings structured into task-centered episodes) and supervised fine-tuning combining human demonstrations, real-robot trajectories, and simulated experience; across 28 embodied spatial-intelligence and planning benchmarks the 8B version reaches an overall average of 72.
medRxiv

Privacy-Aware Distillation of Large Language Models for Enhanced Multimorbidity Scoring

This study introduces and evaluates a privacy-preserving knowledge distillation framework in which CTGAN-generated synthetic cohorts matching UK Biobank distributions are used to elicit multimorbidity scores from three teacher LLMs (GPT-4o, Gemini, DeepSeek) under zero-shot prompting, and compact student models (CoLLMs) are then trained to mimic those scores, enabling application to real UK Biobank data (N = 439,221) for multimorbidity scoring without exposing patient-level data to third-party APIs, with evaluation against the Charlson (CCI) and Elixhauser (ECI) indices via survival analysis, genome-wide association studies, and polygenic risk score associations.
Plastic and reconstructive surgery

Detecting Manipulated Online Rhinoplasty Images Using an Artificial Intelligence Facial Authenticity Localization (FAL) Model

This study provides the first large-scale application of an AI-based Facial Authenticity Localization (FAL) model to aesthetic-surgery media, analyzing 600 consecutive postoperative rhinoplasty photographs from RealSelf.com's public gallery and finding suspected digital manipulation in 19.5% (117/600; 95% CI, 16.3-22.7%), with a 1.5% false-positive rate (3/200; 95% CI, 0.51-4.32%) on 200 presumed-unedited clinical photographs and 100% sensitivity (50/50; 95% CI, 93.0-100%) on 50 Facetune-generated geometric warps.
PloS one

MEF-TBN: A Three-Branch Network for Multi-Exposure Image Fusion with Feature Extraction and Color Enhancement

The work proposes MEF-TBN, a three-branch network operating in YCbCr space that uses a context aggregation attention network (CAAN) branch for multi-scale local features, a transformer branch with multi-head self-attention for global long-range dependencies, and a color enhancement branch that learns the mapping between luminance and chrominance; the two feature branches produce low-resolution weight maps that are refined by a merger module with three merger blocks and upsampled to original resolution via guided filtering for joint upsampling (GFU), and across three MEF datasets against 9 representative methods it achieves the best results on image-feature metrics such as AG, EI and SF and on all human-perception metrics such as QCB, VIF and NIQE, ranks second on MEF-SSIM, and improves ove
medRxiv

Artificial Intelligence-Enhanced Electrocardiography for Detection and Prediction of Hypertrophic Cardiomyopathy across Monogenic and Polygenic Susceptibility

In 1,095 carriers of pathogenic or likely pathogenic sarcomere variants across three international centers, a previously validated AI-ECG model applied to 12-lead ECG images yielded an AUROC of 0.91 for the HCM phenotype at baseline and 0.92 for manifest HCM, a higher AI-ECG score among genotype-positive/phenotype-negative individuals predicted incident HCM during follow-up (unadjusted HR 1.55 per 1-SD; adjusted HR 1.38), and in 57,007 UK Biobank participants the AI-ECG score and a polygenic risk score were independent and additive, with adjusted odds of HCM of 60.2 when both were high.
PLoS biology

AlloPool: a graph neural network framework that infers protein allostery from molecular dynamics simulations

The study presents AlloPool, a graph neural network framework that combines temporal attention with iterative edge pooling to learn minimal, time-evolving residue interaction networks from equilibrium and non-equilibrium molecular dynamics trajectories, reconstructing trajectories at sub-angstrom RMSD across Pin1, the GAIN mechanosensor domain, dopamine D2 and beta-1 adrenergic receptors, the SdrG adhesin, PDZ3, and the Engrailed homeodomain, and using those networks to map allosteric pathways, predict ligand pharmacology, mechanical loading states, and mutation effects.
medRxiv

Decision Support in Publicly Available Patient Information Policies at U.S. Osteopathic Medical Schools: A Vignette-Based Document Analysis

Using a sample of 20 U.S. osteopathic medical schools and eight educational vignettes yielding 160 school-vignette pairs, each assessed in three AI-assisted retrieval-and-evaluation runs and classified as Explicitly Supported, Inferable, Ambiguous, or Not Addressed, the study found that at least one eligible source was retrieved for 159 of 160 pairs (99.4%) yet only 21 pairs (13.1%) were grouped as sufficiently supported, most often for generative AI-assisted reflective writing (7/20 schools, 35%) and in no case for personal cloud notes or official clinical logs.
medRxiv

The Online Health Safety Gap: Consensus Alignment Does Not Imply Safety in Peer-to-Peer Health Narratives

This work identifies and empirically validates the Online Health Safety Gap: across 704 health narratives (699 classified), two scores computed from non-overlapping feature sets, Narrative Truth Distance for epistemic divergence and Narrative Risk Score for health risk potential, share only 4.9% of their variance (r=0.222, p<0.001), so 39.6% of narratives fall in the two off-diagonal quadrants that single-axis systems mishandle by construction (aligned-but-risky 25.2%, divergent-but-safe 14.4%), and on an expert-labeled misinformation benchmark of 437 posts containing 127 misinformation instances the highest discrimination is Youden J of 0.349, motivating a Classification Quadrant and a shift from fact-centric evaluation to risk-aware assessment.
medRxiv

MES: A Multi-Agent Evidence Synthesis System for Medical Decision-Making

This work presents MES, a multi-agent framework of six specialized agents (Supervisor, LiteratureMiner, RWD-Analyst, KG-Specialist, Statistician, SafeChecker) that integrates published literature and trial evidence, generates question-specific real-world evidence from real-world clinical data, and queries biomedical knowledge graphs while preserving source traceability; evaluated in two clinical use cases (an Alzheimer disease medication question with no matching literature and a septic shock beta-blocker question with inconsistent randomized evidence) and across 144 clinical queries spanning six evidence-based medicine categories, it produced structured reports that preserved source traceability, identified cross-source disagreement and communicated uncertainty, and its SafeChecker improv
PLOS global public health

Reframing the health workforce as a hybrid human–AI capability asset: the proposal of Workforce Science

This Opinion article proposes Workforce Science, arguing that amid population ageing, epidemiological transition, widening inequities and the rapid embedding of artificial intelligence, the health and care workforce should be reframed from a supply question of how many workers to educate and employ into a question of how hybrid human–AI capability, generated jointly by people, teams and intelligent technologies, is measured, analysed and stewarded; it offers a construct of effective capability (C*) as a function of education financing, education quality, labour, the human–AI dynamic, governance capacity and the conditions, rights and realities of work, together with a five-layer path of definition, measurement, analysis, governance and stewardship.
PloS one

Explainable AI for Sentiment Analysis of HMPV Using XLNet and SHAP

The study scraped 15,300 HMPV-related comments from YouTube news channels in 2024-2025, retained 9,758 after preprocessing, generated weak sentiment labels with VADER, compared six transformer models (ELECTRA, RoBERTa, ALBERT, DistilBERT, XLNet, BERT), found XLNet best at 93.50% accuracy, and used SHAP's PartitionExplainer to produce word-level attributions on correctly classified samples, showing words such as "flu" and "fear" driving negative predictions, "save", "help" and "mad" associated with neutral, and "falling", "save", "god" and "america" associated with positive.
medRxiv

Simulating Community Health Behavior with LLM Synthetic Populations: A Cross-State Evaluation of LLMPopSim

The study introduces LLMPopSim, a generative population simulation framework that integrates U.S. Census and CDC data to construct synthetic individuals and uses an LLM to simulate individual health behaviors whose aggregate outcomes can be evaluated at the community level; developed with historical Hawaiʻi data and evaluated for temporal and geographic generalizability using held-out 2022 cohorts from Hawaiʻi and New York State with colorectal cancer screening and mammography as proof-of-concept behaviors, the four state-outcome evaluations showed mean absolute error of 3.5 to 15.0 percentage points and correlations between simulated and observed ZCTA-level prevalence of 0.26 to 0.69.
medRxiv

Fully Automated Abstraction of Longitudinal Breast Oncology Records with Off-The-Shelf Large Language Models

The study developed a HIPAA-compliant open-source pipeline in which off-the-shelf commercial large language models, without fine-tuning, abstracted variables from unnormalized, unlabeled, and unedited clinical notes, pathology reports, medication administration records, and demographics for 100 complex breast cancer patients (median chart over 3,100 pages, median 6.5 years of follow-up, median 7 lines of therapy), achieving high concordance with an expert oncologist for recurrence status (99%), germline BRCA1/2 pathogenic variants (100%), hormone receptor status (99%), HER2 status (96%), clinical stage (91%), PIK3CA mutation status (91%), and ESR1 mutation status (90%), approaching inter-oncologist variability for anti-cancer drug extraction, while exact therapy-line reconstruction remaine
bioRxiv

EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models

EvSpark is a speculative decoding system for hybrid convolutional-recurrent-attention DNA foundation models such as Evo2: it verifies draft blocks in parallel and restores all three classes of inference state by selecting retained intermediate states without replay, reaching 2.96x on 43 real-sequence prompts and 3.27x including five synthetic controls on an Evo2 7B 48-prompt benchmark with three training seeds, retaining 1.84x-2.43x at 262k context, and yielding 2.18x-2.46x on real sequences and 2.51x-2.78x on the full suite for 20B and 40B targets.
EuroIntervention

Radial wall strain for residual risk stratification after percutaneous coronary intervention

Using the TARGET All Comers randomised trial of 1,551 patients, this study developed an artificial intelligence algorithm for real-time automated assessment of radial wall strain (RWS) from routine angiography and evaluated its prognostic value in 1,384 non-target vessels from 802 patients, finding that baseline maximal RWS (RWSmax) ≥13% independently predicted the 5-year non-target vessel-oriented composite endpoint (NT-VOCE) with an adjusted hazard ratio of 4.82 (95% CI 3.14-7.40) and an adjusted AUC of 0.73, with the highest predictive accuracy for non-target vessel revascularisation (adjusted AUC 0.92) and even greater performance in the high-quality angiographic image subgroup (adjusted HR 6.89).
arXiv

Foresight Without Trajectories: Guiding 3D Diffusion Policies with a Latent Movement Trend

The work introduces Movement Trend Guidance: a compact latent of interaction evolution is learned from a short history of point clouds and robot states, supervised during training only by sparse future gripper states, and at inference retained alone as future-oriented conditioning, consistently improving 3D diffusion policies on RoboTwin2.0, LIBERO-40, DexArt and five real-robot tasks while adding only 3.52% more parameters than DP3.
arXiv

MLLM Hallucinations Arise When Information Distribution Drifts Inside Synergy Heads

The work proposes HEAL, which applies causal noise intervention to multi-head outputs and counterfactual Difference-in-Differences to categorize attention heads into redundant, visual, language, and synergy types, finds that hallucinations occur when the visual-language information distribution inside synergy heads drifts away from a healthy equilibrium rather than correlating strongly with the quantity or strength of modality-specific heads, and accordingly injects dynamic calibration factors into the value vectors of synergy heads at inference time, reducing hallucinations across multiple MLLMs and benchmarks.
arXiv

Paint-Anything: Unified Any-Color Control for Image Generation and Editing

Paint-Anything introduces a shared hex-prompt interface that achieves any 24-bit hex target-color control for both image generation and editing through object-level color supervision, builds the Paint-500K data pipeline and the ACBench benchmark, and on FLUX.2-4B improves ACBench-T2I and ACBench-Edit scores by 85.3% and 28.3% respectively relative to the base model while attaining the highest average CompColor score among compared methods.
arXiv

EvoOntology: A Self-Evolving Ontology Layer for Data Agents

The work introduces EvoOntology, which encapsulates an ontology as an MCP server with schema, content, and tool layers, builds an initial ontology via a builder agent issuing probe queries, and continuously refines it through attribution-guided typed edits admitted only after backbone-conditional paired evaluation, consistently outperforming ReAct baselines and static semantic-layer baselines on three data-agent benchmarks with six LLM backbones.
arXiv

Letting the Sparsity Coefficient Learn Its Own Compression: Training-Adaptive Convolutional Sparse Coding Through an Information Bottleneck Lens

This work turns the sparsity coefficient of convolutional sparse coding from a manually fixed hyperparameter into a differentiable variable jointly learned with network parameters through unfolded FISTA iterations, interprets that coefficient within an information bottleneck framework as the controller of the compression-retention trade-off, achieves competitive clean-data recognition on CIFAR-10/100 and ImageNet-1K with greatly improved robustness under various input perturbations, and introduces a label-free post-training adaptation strategy that adjusts compression strength using a small set of unlabeled corrupted samples while keeping the main network parameters frozen.
arXiv

Writing Agent Skills as Graphs: GraphSkillEvo Evolves Reusable Procedural Knowledge

The work represents LLM agent skills as graph-structured natural-language artifacts, where each node is an execution step with its operational guidance and directed edges encode context-dependent transitions between steps, and builds GraphSkillEvo, a population-based evolutionary framework with global-guidance mutation, graph-structure mutation, global-guidance crossover, and graph-structure crossover that searches this structured skill space, outperforming the skill-optimization baseline SkillOpt on average across five agent benchmarks, two LLMs, and two execution settings (+4.01% on GPT-5.4-nano and +1.76% on GPT-5.4) while consuming fewer total optimization tokens.
arXiv

MintAct: A Unified Visual Agent for Digital Environments

MintAct introduces a family of vision-language models at 2B, 4B, and 8B scales that unify UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use in a single set of weights through a shared screenshot-and-pixel observation space, prompt-conditioned per-domain action sets, and balanced cross-domain mixing, backed by asynchronous RL infrastructure hosting hundreds of concurrent instances, reaching 48.9 on OSWorld-Verified, 39.1 on Online-Mind2Web, and 67.0 on AndroidWorld while matching or exceeding size-matched per-domain specialists.
arXiv

A Refreshable, Near-Domain-Controlled Audio Benchmark for Telecom Fraud Detection: TeleAntiFraud 2.0

The work presents TeleAntiFraud 2.0, a refreshable Chinese audio benchmark for telecom fraud detection organized as monthly frozen snapshots, in which mixed-tree generation makes fraud and lawful near-domain calls share context and diverge only at label-bearing actions, and it reports that three text classifiers reach perfect Macro-F1 against unrelated or ordinary negatives but drop to 0.65-0.68 with near-domain sibling negatives, while full-set audio and ASR+LLM evaluations expose class-prior shortcuts, prediction collapse, and snapshot sensitivity.
arXiv

MoME: Turning Sparse Lookup into Context-Aware Memory Mixtures

The work introduces Mixture of Memory Embeddings (MoME), which replaces each token's single memory row with multiple memory slots and uses a learned gate over the hidden state to sparsely choose which slots to read at each position, giving memory retrieval a context-aware form while keeping cheap token-indexed lookup; in controlled pretraining across nanochat, Llama-3/MobileLLM, and Qwen3 backbones, MoME improves over Value Embedding, Bigram, and STEM baselines in iso-parameter and iso-training-FLOP settings, shows a more favorable memory-size scaling trend at sub-billion scale, remains efficient in training and inference, and routing analyses on polysemous tokens suggest the same surface token is dispatched to distinct memory slots under different senses.
arXiv

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Starting from Llama 3.1 8B, the study first fine-tunes a reviewer on official ICLR reviews from 2018–2023, then trains four successor models on ICLR 2024 data with synthetic reviews generated by that reviewer mixed at 0%, 33%, 66%, and 100%, finding that higher synthetic exposure compresses rating distributions and monotonically reduces same-paper and corpus-level semantic diversity (about 11% and 5%), a pattern the authors call scientific-judgment collapse, and introduces TrustReviewer, which mitigates this tendency through training-time corpus curation and test-time paired activation steering.
arXiv

Treating Refinement Itself as the Editing Interface: How RefineEdit Edits Images Without Training in a Generative Refinement Network

The work introduces RefineEdit, a training-free prompt-to-prompt image editing framework that couples edit localization with content generation inside the global refinement of binary image codes in a Generative Refinement Network (GRN): it branches an editing branch from an intermediate source state, uses the signed differences in the two branches' probabilities for the same source-sampled bits to select editable positions and bits, copies the evolving source state at all remaining bits, and stabilizes decisions across steps with adaptive spatial freezing and finite bit locking, achieving the best background-preservation scores (PSNR, LPIPS, MSE, SSIM) and the highest whole-image and edited-region CLIP scores among the evaluated methods across nine editing categories of PIE-Bench.
arXiv

Growing Training Grounds from Code Itself: How CodeMidas Turns Implemented Functionality into RL Environments for Coding Agents

CodeMidas presents an agentic pipeline that uses source code as its only task-specific input to turn implemented functionality in open-source codebases into executable, verifiable coding RL environments, yielding 5,545 training tasks from 3,185 codebases across 23 programming languages and 15 technical domains, and training MiMo-V2.5 with GRPO improves all five external benchmarks, including DeepSWE pass rate from 10.0% to 21.7%, ProgramBench Almost Solved from 4.5 to 21.5, and Terminal-Bench v2.1 from 63.7% to 72.2%.
arXiv

DeformSmith: Turning Text or a Single Image into Graspable Deformable Assets via a Physics Harness

DeformSmith presents a hierarchical agentic framework that, from text or a single image, progressively builds geometry, a physical model, material behavior, and robot interaction, using a shared physics-grounded harness to evaluate and revise candidates through simulation probes and robot pick-and-place feedback; across 39 cases its assets score higher on visual quality and physical plausibility than baselines including PhysGen3D, PhysGM, and PhysX-Omni, while also producing replayable manipulation data.
arXiv

OmniVBench and Omni-R2V: A Shared Foundation for Evaluating and Training Omni Reference-to-Video Generation

This work introduces the OmniVBench benchmark and the Omni-R2V dataset, evaluating omni reference-to-video generation with 7 task families, 18 fine-grained tasks, 813 evaluation cases and 12,172 factor-grounded checklist items, and releasing 340K processed training samples; evaluation of several open- and closed-source models reveals clear performance gaps across task families and evaluation dimensions.
arXiv

Separating Teacher Self-Deviation from the Distillation Signal: Calibrated On-Policy Distillation

The work argues that the token-level teacher–student discrepancy used by on-policy distillation (OPD) mixes in context-induced teacher-side variation, termed Teacher Self-Deviation (TSD), and introduces Cal-OPD, which estimates the teacher's self-deviation region via positive and negative privileged interventions and keeps only the residual beyond that region as the optimization signal, consistently outperforming standard OPD and its variants on mathematical reasoning benchmarks while retaining only about 52–65% of the original discrepancy.
arXiv

Retention-Constrained Post-Training Quantization of Cellpose–SAM: An Auditable Compression Protocol for Stem Cell Microscopy

The work proposes a pre-specified retention protocol and applies it to several post-training quantization schemes for Cellpose–SAM on a stratified 176-field public panel spanning BBBC038 nuclei, BBBC039 U2OS fluorescence, and NIST iPSC images: weight-only W8A16 preserves instance F1 across all modalities, a sensitivity-guided mixed W4/W8 scheme with four INT8 exception operators reduces weight storage from 1162.07 MiB to 171.86 MiB with no observed catastrophic failures, while ternary W2A16-G64 compresses to 96.21 MiB but fails catastrophically on 169 of 176 fields, showing that compression should be judged by modality-stratified downstream retention rather than a single accuracy number.
arXiv

RecreationWorld: Stitching Interface Operation and Code Writing into One Long-Horizon Trajectory via Recreation

This work introduces RecreationWorld, a framework, and RecreationBench, a benchmark: given a running reference application, an agent must discover its behavior and deliver a buildable, runnable recreation, with reproducible environments on Ubuntu, macOS, Windows, Android, and Web plus a unified GUI-and-coding harness, automatic scoring by reference-derived hidden programmatic and visual assertions, and 35,000 filtered recreation trajectories for training; evaluation shows GPT-6 Astra leading at 58.1% overall while passing all programmatic tests on just 2.8% of tasks, and recreations reproducing static interface structure more reliably than interactions and computed outputs.
arXiv

Structured Skill Optimization Under Frozen Weights: A New Path for Audio Anti-Fraud Detection

Without changing any parameters of the audio-language model, FRAUDSkill optimizes an external layer of skill programs, routing policies, and decision rules and combines structured output control with validation-guided multi-path inference, raising Macro-F1 on the TeleAntiFraud benchmark from 41.54% for the shared frozen-model baseline to 73.50% while cutting the invalid-output rate from 36.04% to 1.94%.
arXiv

Synthesizing Verifiable Skills from Code at Scale: A Column on Code2Skill and CodeSkillBank

The work introduces Code2Skill, a fully automated pipeline that abstracts source units from 19,769 actively maintained GitHub repositories into atomic-operation, composite-workflow, and recurring-pattern skill records, verifies each record through source-body-blind reconstruction and source-aware comparison, and builds CodeSkillBank with 1,006,822 accepted records, improving 57 of 72 protocol-matched evaluations by 11.7% on average and outperforming trajectory-derived skill banks on all seven shared benchmarks.
arXiv

IntBMoE: Decoupling Expert Participation, Execution, and Materialization via Block-Level Conditioning

The work proposes IntBMoE, a block-conditioned Mixture-of-Experts architecture in which a small learned codebook defines reusable multi-layer blocks, a shared hypernetwork merges all expert bases of each layer into one composed expert, and a router executes only a few blocks per token, thereby keeping pool-wide participation while retaining sparse execution and bounded parameter materialization; it reaches 73.76% Top-1 and 91.48% Top-5 on ImageNet-1K, outperforms the compared sparse and dense MoE baselines on MiniPile language modeling and IntTravel sequential recommendation, and is fully deployed in AMap's generative recommendation system with a 2.4% relative UVCTR gain in online A/B testing.