AI Core
903 items
Study combines Saussurean semiotics, brainstorming and AIGC to generate 20 leather bag splicing designs, with B27 scoring highest satisfaction
Taking splicing methods in leather bag design as its object, the study first used Saussurean semiotics to analyze the symbiotic relationship between "signifiers" such as material, technique and element splicing and their corresponding "signifieds" of function, aesthetics and connotation, then determined target demand and style positioning through market research, constructed signifier-signified relationships with brainstorming and structured them into AIGC prompts to batch-generate 20 series of leather bag schemes, and finally ranked and selected the best through a CSAT satisfaction survey; results show that the individual alpha values of the 20 bags all exceed 0.80, KMO=0.8664 with a significant Bartlett sphericity test (P<0.001), average overall satisfaction ranges from 3.34 to 3.
Legate-Yang and Massenkoff build a robot exposure index: robots can already do 74% of US physical tasks but are cost-competitive for just 0.3% of work
Using Claude to score roughly 19,000 O*NET task descriptions across about 900 occupations by how controlled an environment a robot needs, the authors build a robot exposure index and find that today's robots can perform 74% of US physical tasks (34% of working hours) but are cost-competitive for only 0.3% of work, while a backtest since 1977 shows jobs with higher exposure later saw wage and employment declines.
DC-SAE splits image tokenization into semantic and pixel branches, hitting 32x compression with 3.37 gFID and 29.79 PSNR on ImageNet and a 4.41x speedup over DC-AE
Peking University and Singapore University of Technology and Design propose DC-SAE, a decoupled compact semantic autoencoder that pairs a frozen semantic encoder for a generation-friendly compact latent with a trainable pixel encoder for low-level detail, achieving 29.79 PSNR and 3.37 gFID on ImageNet at 32x spatial compression, improving over DC-AE by 13.5% in PSNR and 54.9% in gFID with a 4.41x speedup, while a B-parameter DiT reaches 0.84 on GenEval and 86.007 on DPG-Bench.
MILO co-evolves agent harnesses and their search strategy, reaching 86.1% on Terminal-Bench 2.1 above the official leaderboard top and tightening three mathematical bounds
MILO (Meta-evolutionary Island Orchestration) co-evolves an agent harness together with the strategy that discovers it: a hierarchical island-based lineage memory treats rejected mutations as negative evidence, per-island mutator agents rewrite complete harnesses, and an orchestrator adapts the search through lineage grafting, speciation, mutator reassignment, and curriculum revision; across Terminal-Bench 2.1, PaperBench, and DeepSWE it improves resolution over its initial harness by 12.0%, 28.3%, and 10.3% with Opus 4.8 (versus best prior-search gains of 4.5%, 18.3%, and 0%), reaches 86.1±2.0% on Terminal-Bench 2.1 (official leaderboard top 83.8±2.
PivotOPD concentrates distillation on pivotal mistakes and the turns after them, lifting Qwen3-1.7B by 5.5 points over the strongest baseline on ALFWorld
Analyzing failed ALFWorld rollouts of three Qwen3 models (8B–235B), the work finds that more than half of failures contain an early and usually recoverable pivotal mistake, and proposes PivotOPD: a teacher model locates pivotal turns and names a gold action plus recovery actions at the next few turns, preventive distillation uses reverse KL on the committed mistake while recovery distillation uses forward KL on self-teacher responses, both folded into a single PPO update with group-based RL, achieving the best average performance on ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students and transferring to a Nemotron-3.5 student on SWE-Bench Verified (+3.2% resolve rate).
SkillGym turns 184k community skills into verifiable training environments, and its 9B fine-tuned model beats a 397B untrained model on two benchmarks
SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.
TERRA reconstructs terrain from kinematics alone, letting a muscle-actuated body complete stairs, ramps and seats
TERRA presents an end-to-end pipeline that, from scene-less kinematic trajectories alone, combines terrain priors, estimated contacts and negative free-space evidence to recover task-relevant support geometry, adds anatomical, tendon-continuity and contact constraints during retargeting, and uses motion-terrain pairs from five datasets to train a single muscle-actuated control policy, improving terrain accuracy, sharply reducing anatomical and interaction violations, and achieving the highest completion rate over supported terrain families across reconstruction, retargeting and held-out tracking benchmarks.
Cross-meeting speaker attribution: ThyVoice leads all evaluated commercial cascades at a 47.13 mean SI-cpWER, while per-meeting leader ElevenLabs gains about 30 percentage points of attribution error on CHiME-6
The work introduces SI-cpWER, a metric that scores cpWER under one corpus-global speaker-ID map, and evaluates five commercial diarize-then-identify cascades, two open baselines, and the end-to-end reference system ThyVoice on CHiME-8 NOTSOFAR (clean and noise-augmented) plus CHiME-6, finding that requiring one persistent identity across meetings changes the commercial ranking: ThyVoice records lower SI-cpWER than every evaluated commercial cascade in all three conditions, with a full-panel mean of 47.13 versus 54.75 for the next system.
Fudan team rewards Chinese glyph structure via IDS decomposition, letting Qwen-Image lead on both structural quality and semantic alignment on LongText and GenTextEval
The work introduces IDSpect: it recursively decomposes target Chinese characters into Ideographic Description Sequence (IDS) tokens using the Unicode 16.0 BabelStone lexicon, trains an expert IDS recognizer built on an SVTRv2 backbone with an NRTR-style decoder to transcribe rendered text regions directly into IDS sequences, and fuses a token-level F1 with globally unique token credit against a whole-character semantic reward, so that GRPO post-training of Qwen-Image improves structural quality and semantic alignment on LongText and GenTextEval without changing the image generator or adding inference-time cost.
Page 3 · showing 10